Barcoded XTEN polypeptides and compositions thereof, and methods for making and using the same
XTEN polypeptides with non-overlapping motifs and barcode fragments enable precise identification and quantification of protein mixtures, addressing the limitations of existing methods and enhancing safety and efficacy by distinguishing full-length from truncated forms.
Patent Information
- Application Number
- JP2025130787
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-11-13
- Filing Date
- 2025-08-05
- Publication Date
- 2025-12-09
AI Technical Summary
Existing methods for identifying and quantifying protein-based prodrugs suffer from limited sensitivity, efficiency, and effectiveness in detecting and distinguishing full-length polypeptides from their size variants, which can impact safety and efficacy due to potential cytotoxicity, immunogenicity, and undesirable pharmacokinetic properties.
Development of XTEN polypeptides with non-overlapping sequence motifs and barcode fragments that are releasable upon protease digestion, allowing for precise identification and quantification of polypeptide mixtures through unique barcode fragments and reference fragments.
Enhances the ability to assess the integrity and relative amounts of full-length and truncated polypeptides, reducing immunogenicity and improving therapeutic efficacy by ensuring accurate dosage and minimizing unintended effects.
Smart Images

Figure 2025179058000001_ABST
Abstract
Description
[Background technology]
[0001] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy, created on November 6, 2020, is named 20-1761-WO_Sequence_Listing_ST25.txt and is 1494 bytes in size.
[0002] Polypeptides can be produced in a manner that results in a mixture of polypeptides. A mixture of polypeptides can often contain full-length polypeptides along with their size variants (e.g., truncations). The presence of variants whose size differs from the desired full-length product can affect the biological behavior of the polypeptide drug substance, potentially impacting its safety and / or efficacy. For example, protein-based prodrugs for cancer therapy can be engineered using a tumor-targeting activation mechanism. More specifically, full-length therapeutic proteins can be produced and administered in an inactive (non-cytotoxic) prodrug form, which is converted to an active drug by preferential removal of a portion of the prodrug polypeptide in the intended biological site (e.g., tumor). Truncation variants of the full-length construct can lose protective sequences and become cytotoxic (active), thereby "contaminating" the prodrug composition and generating a mixture with components that are unintentionally active outside the intended biological site. In some cases, such shorter-length variants may pose a greater risk of immunogenicity, have lower selective toxicity to tumor cells, or exhibit less desirable pharmacokinetic properties (e.g., resulting in a narrower therapeutic window) than the full-length protein, or may have adverse unintended effects in recipients outside the intended site (e.g., in healthy tissue). As a result, the detection and quantification of protein structural variations can be important for evaluating the biological properties (e.g., clinical safety and pharmacological efficacy) of biological therapeutics and for developing new biological therapeutics (e.g., with increased efficacy and reduced side effects). Existing techniques and methods for identifying and quantifying the amount of "contaminating" cleavage products may include one or more drawbacks, such as limited sensitivity, ease, efficiency, or effectiveness. Summary of the Invention
[0003] Disclosed herein are polypeptides comprising extended recombinant polypeptides (XTENs) consisting of multiple non-overlapping sequence motifs. In the XTEN polypeptides of the invention, the multiple non-overlapping sequence motifs comprise a set of non-overlapping sequence motifs, each of which is repeated at least twice in the XTEN polypeptide and is a unique non-overlapping sequence motif that occurs only once within the XTEN polypeptide, and the polypeptide further comprises a first barcode fragment releasable from the polypeptide upon protease digestion. In the foregoing embodiment, the first barcode fragment is a portion of the XTEN that comprises at least a portion of a sequence motif that occurs only once within the XTEN and differs in sequence and molecular weight from all other peptide fragments releasable from the polypeptide upon complete protease digestion of the polypeptide. Furthermore, in embodiments of the XTENs of the invention provided herein, the barcode fragment does not comprise the N- or C-terminal amino acids of the polypeptide. As further disclosed herein, the XTEN polypeptides of the invention are characterized as comprising at least 150 amino acids in length, more specifically, between 150 and 3000 amino acids in length. The amino acid residues comprising the XTEN polypeptides of the invention are characterized such that at least 90% of these residues are glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), or proline (P), and the XTEN polypeptides are characterized such that at least 90% of these residues are glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), or proline (P). The XTEN polypeptides provided herein further comprise non-overlapping sequence motifs that are sequences of 9 to 14 amino acids in length, and within each of said non-overlapping motifs, the sequence of the G, A, S, T, E, or P amino acids is substantially randomized with respect to any other non-overlapping sequence motif that comprises the XTEN polypeptide.
[0004] In some embodiments, a barcode fragment does not contain a glutamic acid immediately adjacent to another glutamic acid in the XTEN. In some embodiments, a barcode fragment has a glutamic acid at its C-terminus. In some embodiments, a barcode fragment has an N-terminal amino acid that immediately precedes a glutamic acid residue. In some embodiments, the glutamic acid residue preceding the N-terminal amino acid is not immediately adjacent to another glutamic acid residue. In some embodiments, a barcode fragment does not contain a glutamic acid residue at a position other than the C-terminus of the barcode fragment, unless the glutamic acid is immediately followed by a proline. In some embodiments, a barcode fragment is located 10 to 150 amino acids from either the N-terminus of the polypeptide or the C-terminus of the polypeptide.
[0005] In some embodiments, the sequence motifs of the set of non-overlapping sequence motifs are identified herein by SEQ ID NOs: 182-203 and 1715-1722. In some embodiments, the sequence motifs of the set of non-overlapping sequence motifs are identified herein by SEQ ID NOs: 186-189. In some embodiments, the set of non-overlapping sequence motifs includes at least two, at least three, or all four of the sequence motifs of SEQ ID NOs: 186-189.
[0006] In certain embodiments, a polypeptide provided herein comprises an XTEN polypeptide disclosed herein, wherein the barcode fragment does not include the N-terminal amino acid or the C-terminal amino acid of the polypeptide, does not include a glutamic acid immediately adjacent to another glutamic acid in the XTEN, has a glutamic acid at its C-terminus, has an N-terminal amino acid that immediately precedes the glutamic acid residue, and is located 10 amino acids to 125 amino acids from either the N-terminus of the polypeptide or the C-terminus of the polypeptide.
[0007] In some of these embodiments, the glutamic acid residue preceding the N-terminal amino acid is not immediately adjacent to another glutamic acid residue, and in some of these embodiments, the barcode fragment does not contain a glutamic acid residue at any position other than the C-terminus of the barcode fragment, unless the glutamic acid is immediately followed by a proline.
[0008] In some embodiments, the XTEN polypeptides provided herein comprise a plurality of non-overlapping sequence motifs, each of which is repeated at least twice in the XTEN polypeptide and is 9-14 amino acids in length. In some embodiments, the sequence motifs of the set of non-overlapping sequence motifs are identified herein by SEQ ID NOs: 182-203 and 1715-1722. In some embodiments, the sequence motifs of the set of non-overlapping sequence motifs are identified herein by SEQ ID NOs: 186-189. In some embodiments, the set of non-overlapping sequence motifs comprises at least two, at least three, or all four of the sequence motifs of SEQ ID NOs: 186-189. In some embodiments, at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid residues of the XTEN polypeptide are a combination of glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), or proline (P), and the XTEN polypeptide comprises at least four of these amino acids (G, A, S, T, E, or E). In some embodiments, the XTEN is 150-3000 amino acids in length. In some embodiments, the XTEN is 150-1000 amino acids in length. In some embodiments, the XTEN is 150-3000 amino acids in length. The polypeptide can be cleaved by a protease that cleaves C-terminal to a glutamic acid residue that is not followed by a proline. In certain embodiments, the protease is a Glu-C protease.
[0009] In some embodiments of the XTEN polypeptides provided herein, the barcode fragment is located within 200, 150, 100, or 50 amino acids of the N-terminus of the polypeptide. In some embodiments, the barcode fragment is located between 10-200, 30-200, 40-150, or 50-100 amino acids from the N-terminus of the protein. In some embodiments, the barcode fragment is located within 200, 150, 100, or 50 amino acids of the C-terminus of the polypeptide. In some embodiments, the barcode fragment is located between 10-200, 30-200, 40-150, or 50-100 amino acids from the C-terminus of the protein. In some embodiments, the barcode fragment is at least 4 amino acids in length. In some embodiments, the barcode fragment is 4-20, 5-15, 6-12, or 7-10 amino acids in length. In some embodiments, the barcode fragments are identified herein by SEQ ID NOs: 8020-8030 (BAR001-BAR011).
[0010] In some embodiments, the polypeptide further comprises a second barcode fragment, the second barcode fragment being part of an XTEN that differs in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide upon complete digestion of the polypeptide by a protease. In some embodiments, the polypeptide further comprises a third barcode fragment, the third barcode fragment being part of an XTEN that differs in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide upon complete digestion of the polypeptide by a protease.
[0011] In some embodiments, the XTEN has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to a sequence identified herein by SEQ ID NOs: 8001-8019. In some embodiments, the XTEN is at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids in length.
[0012] In some embodiments, the polypeptide further comprises a biologically active polypeptide linked to an XTEN polypeptide (BPXTEN). In some embodiments, the XTEN polypeptide is linked to the biologically active polypeptide at the amino or carboxyl terminus of the XTEN. In either configuration, the barcode fragment is located within a region of the XTEN that spans 5% to 50%, 7% to 40%, or 10% to 30% of the length of the XTEN, as measured from the amino or carboxyl terminus linked to the biologically active polypeptide.
[0013] In some embodiments, the BPXTEN polypeptide further comprises one or more reference fragments releasable from the polypeptide upon digestion with a protease, each of the one or more reference fragments comprising a biologically active portion of the polypeptide. In some embodiments, the one or more reference fragments are a single reference fragment that differs in sequence and molecular weight from all other peptide fragments releasable from the polypeptide upon digestion of the polypeptide with a protease. In some embodiments, the reference fragment comprises a peptide whose presence in a polypeptide mixture indicates its presence or integrity (i.e., that has not been proteolytically degraded or proteolytically cleaved).
[0014] In some embodiments, the BPXTEN polypeptide is biologically active with XTEN. The polypeptide further comprises a first release segment (RS1) located between the polypeptides. In some embodiments, RS1 comprises an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a sequence identified herein by any one of the sequences in Tables 4a-4h. In some embodiments, the biologically active polypeptide is identified herein by any one or combination of the sequences in Tables 4a-4h and 8a-8b.
[0015] In some embodiments, the BPXTEN polypeptide advantageously has a terminal half-life that is at least 2-fold longer compared to a biologically active polypeptide that is not linked to XTEN.
[0016] In some embodiments, BPXTEN polypeptides are advantageously less immunogenic than biologically active polypeptides that are not linked to XTEN, and immunogenicity can be confirmed by measuring the production of IgG antibodies that selectively bind to the biologically active polypeptide after administration of equivalent doses to humans or animals.
[0017] In some embodiments, the BPXTEN polypeptide exhibits an apparent molecular weight factor under physiological conditions of greater than about 6.
[0018] In some embodiments, the BPXTEN polypeptide further comprises a second XTEN polypeptide, wherein the second XTEN polypeptide comprises an amino acid sequence having some of the characteristics described above and throughout this disclosure for the first XTEN component of these embodiments of BPXTEN, wherein the first XTEN polypeptide is located at the N-terminus of the biologically active polypeptide and the second XTEN polypeptide is located at the C-terminus of the biologically active polypeptide. In some embodiments, the second XTEN polypeptide comprises an amino acid sequence that differs from the amino acid sequence of the first XTEN that comprises these embodiments of BPXTEN. In certain embodiments, the amino acid sequence of the second XTEN polypeptide is longer than the amino acid sequence of the first XTEN polypeptide.
[0019] In some embodiments, the BPXTEN polypeptide further comprises a second release segment (RS2) located between the biologically active polypeptide and the second XTEN polypeptide. In some embodiments, RS1 of the first XTEN polypeptide and RS2 of the second XTEN polypeptide are identical in sequence. In some embodiments, RS1 of the first XTEN polypeptide and RS2 of the second XTEN polypeptide are each substrates for cleavage by multiple proteases at one, two, three, or more cleavage sites within each release segment sequence.
[0020] In some of these embodiments, the BPXTEN polypeptide further comprises a barcode fragment that is part of a second XTEN polypeptide and that differs in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide upon complete digestion of the polypeptide by a protease. In some of these embodiments, the additional barcode fragment does not comprise the C-terminal amino acid of the polypeptide. In some of these embodiments, the additional barcode fragment comprises a glutamic acid residue at its C-terminus. In some of these embodiments, the additional barcode fragment of the second XTEN polypeptide is located within 200, 150, 100, or 50 amino acids of the C-terminus of the second XTEN component of the BPXTEN polypeptide. In some of these embodiments, the additional barcode fragment of the second XTEN polypeptide is located between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the C-terminus of the second XTEN component of the BPXTEN polypeptide. In some of these embodiments, the additional barcode fragment is 4 to 20, 5 to 15, 6 to 12, or 7 to 10 amino acids in length. In some of these embodiments, further The barcode fragments are identified herein by SEQ ID NOs: 8020-8030 (BAR001-BAR011).
[0021] In some embodiments, the second XTEN polypeptide further comprises a set of barcode fragments comprising an additional barcode fragment and at least one additional barcode fragment, wherein each barcode fragment of the set of barcode fragments differs in sequence and molecular weight from all other peptide fragments liberable from the BPXTEN polypeptide upon complete digestion of the polypeptide by a protease. In some embodiments, the second XTEN polypeptide is identified by SEQ ID NOs: 8001-8019. In some embodiments, the additional barcode fragment does not comprise a glutamic acid residue immediately adjacent to another glutamic acid residue in the polypeptide.
[0022] In some embodiments, at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid residues in the second XTEN polypeptide are a combination of glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), and proline (P), and the XTEN polypeptide comprises at least four of these amino acids (G, A, S, T, E, or P). In some embodiments, the sum of the total number of amino acids in the first XTEN polypeptide and the total number of amino acids in the second XTEN polypeptide is at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, or at least 800 amino acids. In some embodiments, the second XTEN polypeptide comprises multiple non-overlapping sequence motifs, each of which is repeated at least twice in the second XTEN polypeptide sequence and is 9 to 14 amino acids in length.
[0023] In some embodiments, for the second XTEN polypeptide, the sequence motif of the plurality of non-overlapping sequence motifs is identified herein by SEQ ID NOs: 182-203 and 1715-1722. In some embodiments, the sequence motif of the plurality of non-overlapping sequence motifs is identified herein by SEQ ID NOs: 186-189. In some embodiments, for the second XTEN polypeptide, the plurality of non-overlapping sequence motifs includes at least two, at least three, or all four of the following motifs: SEQ ID NOs: 186-189. In some embodiments, the second XTEN polypeptide is 150-3000 amino acids in length. In some embodiments, the second XTEN polypeptide is 150-1000 amino acids in length. In some embodiments, the second XTEN polypeptide has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to those identified herein by SEQ ID NOs: 8001-8019. In some embodiments, the second XTEN polypeptide is at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids in length.
[0024] In certain embodiments, the BPXTEN polypeptides provided herein comprise a first XTEN polypeptide covalently linked to a first and second biologically active polypeptide in tandem, the first XTEN polypeptide comprising a first RS sequence proximal to the C-terminus of the polypeptide but not including the C-terminus of the polypeptide, and a second XTEN polypeptide covalently linked to the C-terminus of the tandemly linked biologically active polypeptide, the second XTEN polypeptide comprising a second RS sequence proximal to the N-terminus of the second XTEN polypeptide but not including the N-terminus of the second XTEN polypeptide, wherein the first and second RS sequences can be the same or different. In certain embodiments, the second XTEN polypeptide comprises a longer amino acid sequence than the amino acid sequence of the first XTEN polypeptide. In certain embodiments, the first or second biologically active protein, or both, are specific binding proteins. In certain embodiments, the specific binding protein specifically binds to an antigen or agonist expressed at a desired biological site. In certain embodiments, the desired biological site is a tumor and the antigen is a tumor-specific antigen. In certain embodiments, the first and second biologically active polypeptides are different, including, but not limited to, those with different specific binding affinities.
[0025] Further disclosed herein are nucleic acids comprising a polynucleotide encoding a polypeptide, such as any of the XTEN or BPXTEN polypeptides disclosed herein, or the reverse complement of the foregoing polynucleotides.
[0026] Also disclosed herein are expression vectors comprising any of the polynucleotide sequences disclosed herein and regulatory sequences operably linked to the polynucleotide sequence that regulate the expression or other biological activity of said polynucleotide.
[0027] Disclosed herein are host cells comprising the expression vectors disclosed herein. In some embodiments, the host cells are prokaryotic. In some of these embodiments, the host cells are Escherichia coli. In some alternative embodiments, the host cells are mammalian cells.
[0028] Additionally, disclosed herein are pharmaceutical compositions comprising the polypeptides disclosed herein and one or more pharmaceutically acceptable excipients. In some embodiments, the pharmaceutical compositions are formulated for administration to animals, particularly humans, and such administration may be by any therapeutically effective route of administration. The pharmaceutical compositions disclosed herein may be prepared and used in any formulation known in the art, particularly suited to the route, site, and intended effect on humans or animals.
[0029] Disclosed herein is the use of the polypeptides disclosed herein, particularly BPXTEN polypeptides, in the preparation of medicaments for the treatment of a disease, disorder, or condition in a human or animal. In some embodiments, the disease, disorder, or condition may be cancer.
[0030] Disclosed herein are methods of treating a disease in a human or animal as disclosed above and throughout this disclosure, comprising administering to a human or animal in need thereof one or more therapeutically effective doses of a pharmaceutical composition. In some embodiments, the pharmaceutical composition is administered to the human or animal as one or more therapeutically effective doses administered on a clinically relevant schedule, such as daily, weekly, monthly, or yearly, at clinically relevant doses.
[0031] A mixture comprising multiple polypeptides, specifically XTEN and BPXTEN polypeptides disclosed herein, of various lengths, a first set of polypeptides, each polypeptide of the first set of polypeptides comprising a barcode fragment releasable from the polypeptide by digestion with a protease and having a sequence and molecular weight that differs from the sequences and molecular weights of all other fragments releasable from the first set of polypeptides; and a second set of polypeptides lacking the barcode fragment of the first set of polypeptides; both the first set of polypeptides and the second set of polypeptides each include a reference fragment common to the first set of polypeptides and the second set of polypeptides, the reference fragment being generated by digestion with a protease; Disclosed herein are mixtures in which the ratio of the first set of polypeptides to the polypeptide comprising the reference fragment is greater than 0.7.
[0032] In some embodiments, the ratio of the first set of polypeptides to the polypeptide comprising the reference fragment is greater than 0.8, 0.9, 0.95, or 0.98. In some embodiments, the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides. In some embodiments, the protease is a protease that cleaves C-terminal to glutamic acid residues. In some embodiments, barcode release from polypeptides comprising the first set of polypeptides is promoted by pepsin, elastase, thermolysin, or Glu-C protease. In some embodiments, barcode release is promoted by Glu-C protease. In some embodiments, the protease is not trypsin. In some embodiments, the polypeptides of varying lengths include polypeptides comprising at least one XTEN polypeptide described herein.
[0033] In some embodiments, the first polypeptide set comprises a full-length polypeptide, and the barcode fragment is a portion of the full-length polypeptide. In some embodiments, the full-length polypeptide is any polypeptide disclosed herein, particularly an XTEN or BPXTEN polypeptide. In some embodiments, the barcode fragment does not comprise either the N-terminal amino acid or the C-terminal amino acid of the full-length polypeptide. In some embodiments, the mixture of polypeptides of various lengths differs from each other due to N-terminal truncation, C-terminal truncation, or both N- and C-terminal truncation of the full-length polypeptide.
[0034] 1. A method of assessing the relative amount of a first set of polypeptides in a mixture comprising polypeptides of varying lengths, particularly XTEN and BPXTEN polypeptides disclosed herein, relative to a second set of polypeptides in the mixture, wherein each polypeptide of the first set of polypeptides shares a barcode fragment that occurs once and only once among the polypeptides, each polypeptide of the second set of polypeptides lacks a barcode fragment shared by the polypeptides of the first set, and each individual polypeptide of both the first set of polypeptides and the second set of polypeptides comprises a reference fragment; contacting the mixture with a protease to generate a plurality of proteolytic fragments resulting from cleavage of the first set of polypeptides and the second set of polypeptides, the plurality of proteolytic fragments comprising a plurality of reference fragments and a plurality of barcode fragments; and Disclosed herein are methods that include determining a ratio of the amount of the barcode fragment to the amount of the reference fragment, thereby assessing the relative amount of the first set of polypeptides to the second set of polypeptides.
[0035] In some embodiments, the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides.
[0036] In some embodiments, the protease cleaves polypeptides of various lengths on the C-terminal side of glutamic acid residues that are not followed by proline residues. In some embodiments, the protease is a Glu-C protease. In some embodiments, the protease is not trypsin. In some embodiments, determining the ratio of the amount of barcode fragments to the amount of reference fragments comprises quantifying the barcode fragments and the reference fragments from the mixture after the mixture of polypeptides is contacted with the protease. In some embodiments, the barcode fragments and the reference fragments are identified based on their respective masses. In some embodiments, the barcode fragments and the reference fragments are identified via mass spectrometry. In some embodiments, the barcode fragments and the reference fragments are identified via liquid chromatography-mass spectrometry (LC-MS). In some embodiments, determining the ratio of barcode fragments to the reference fragments comprises isobaric labeling or stable isotope labeling. In some embodiments, determining the ratio of barcode fragments to the reference fragments comprises quantifying the mixture with one or both of an isotopically labeled reference fragment and an isotopically labeled barcode fragment. This includes adding
[0037] In some of these embodiments, the polypeptides of various lengths include full-length polypeptides and their truncated fragments. In some of these embodiments, the mixture of polypeptides of various lengths differs from each other due to N-terminal truncation, C-terminal truncation, or N-terminal and C-terminal truncation of the full-length polypeptide. In some of these embodiments, the ratio of the amount of barcode fragment to the amount of reference fragment is greater than 0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.98, or 0.99.
[0038] Disclosed herein is a mixture comprising a plurality of polypeptides of varying lengths, the mixture comprising a first set of polypeptides, each polypeptide of the first set of polypeptides being releasable from the polypeptide by digestion with a protease and comprising a barcode fragment having a sequence and molecular weight that differs from the sequence and molecular weight of all other fragments releasable from the first set of polypeptides. The aforementioned embodiment also includes a second set of polypeptides lacking the barcode fragments of the first set of polypeptides, and both the first and second sets of polypeptides each comprise a reference fragment that is common to the first and second sets of polypeptides and is releasable by digestion with a protease. In the aforementioned embodiment, the number of reference fragments quantified in the polypeptide mixture after protease digestion is equal to the sum of the number of the first and second sets of polypeptides in the mixture, and the number of barcode fragments quantified in the polypeptide mixture after protease digestion is equal to the number of the first set of polypeptides in the mixture. In the aforementioned embodiment, the first set of polypeptides comprises the reference fragment, and the ratio of the first set of polypeptides to the polypeptides in the mixture comprising the reference fragment is greater than 0.7.
[0039] In some embodiments, the mixture has a ratio of the first set of polypeptides to the polypeptide comprising the reference fragment that is greater than 0.8, 0.9, or 0.95.
[0040] In certain embodiments, the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides, hi alternative embodiments, the reference fragment occurs twice in each polypeptide of the first set of polypeptides and the second set of polypeptides.
[0041] In some embodiments, the first set of polypeptides comprises full-length polypeptides and the barcode fragments are portions of the full-length polypeptides.
[0042] In some embodiments, the full-length polypeptide comprises a polypeptide disclosed herein.
[0043] In certain embodiments, the mixture barcode fragment does not include the N-terminal and C-terminal amino acids of the full-length polypeptide.
[0044] In some embodiments, the mixture contains polypeptides of various lengths that differ from each other due to N-terminal truncations, C-terminal truncations, or both N- and C-terminal truncations of the full-length polypeptide.
[0045] In some embodiments, a reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides, hi alternative embodiments, the number of reference fragments in the first set of polypeptides can be different from the number of reference fragments in the second set of polypeptides, but the number in each polypeptide of each set must be the same.
[0046] In one particular embodiment, each of the reference fragments in the mixture of polypeptides has a sequence and molecular weight that differs from the sequence and molecular weight of every other fragment.
[0047] A mixture comprising a plurality of polypeptides of various lengths, the mixture comprising a first set of polypeptides, each polypeptide of the first set of polypeptides comprising: Disclosed herein is a mixture comprising a barcode fragment releasable from a polypeptide by digestion with a protease, the barcode fragment having a sequence and molecular weight different from the sequences and molecular weights of all other fragments releasable from a first set of polypeptides. The mixture further comprises a second set of polypeptides lacking the barcode fragments of the first set of polypeptides, and both the first set of polypeptides and the second set of polypeptides each comprise a reference fragment common to the first set of polypeptides and the second set of polypeptides, the reference fragment being releasable by digestion with a protease. The ratio of the first set of polypeptides to the polypeptides in the mixture is determined by the equation: [barcode-containing polypeptide] / [(reference peptide-containing polypeptide) × N], where N is the number of occurrences of the reference peptide released from each polypeptide in the mixture, and when the first set of polypeptides contains one reference fragment, The ratio of the first set of polypeptides to the polypeptides in the mixture containing the reference fragment is greater than 0.7.
[0048] In certain embodiments, the ratio of the first set of polypeptides to the polypeptide comprising the reference fragment is greater than 0.8, 0.9, or 0.95.
[0049] In some embodiments, the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides.
[0050] In some embodiments, the reference fragment occurs twice in each polypeptide of the first set of polypeptides and the second set of polypeptides.
[0051] In certain embodiments, the first set of polypeptides comprises full-length polypeptides and the barcode fragments are portions of the full-length polypeptides.
[0052] In some embodiments, the full-length polypeptide comprises a polypeptide disclosed herein. In certain embodiments, the barcode fragment does not include the N-terminal and C-terminal amino acids of the full-length polypeptide.
[0053] In some embodiments, the mixture of polypeptides of various lengths differ from each other due to N-terminal truncations, C-terminal truncations, or N- and C-terminal truncations of the full-length polypeptide.
[0054] In some embodiments, the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides. In further embodiments, the number of reference fragments in the first set of polypeptides can be different from the number of reference fragments in the second set of polypeptides, but the number in each polypeptide of each set must be the same. In some embodiments, each of the reference fragments in the polypeptides of the mixture has a sequence and molecular weight that is different from the sequences and molecular weights of all other fragments.
[0055] A method of detecting sequence integrity of polypeptides comprising a first set of polypeptides in a mixture as disclosed herein, comprising digesting the mixture of polypeptides with a protease that releases barcode fragments and reference fragments from the first set of polypeptides and releases reference fragments from the second set of polypeptides, and determining a ratio of barcode fragments from the first set of polypeptides to reference fragments from the first and second sets of polypeptides. In certain embodiments, the sequence integrity of the polypeptides of the first set of polypeptides is detected by comparing the ratio of the fragments to an expected ratio of the fragments based on the number of barcode fragments and reference fragments in the polypeptides comprising the first and second sets of polypeptides.
[0056] The methods contemplated herein are readily amenable to qualitative and quantitative analysis of polypeptides containing barcodes and / or reference fragments, for example, using LC / MS. In one particular embodiment, the LC / MS is quantitative, detecting isotopically distinguishable amounts of barcode fragments, reference fragments, or both. In an exemplary such method, a known amount of a "standard" is added to a mixture of polypeptides to facilitate such analysis. For example, such a standard may include an isotopically labeled version of the aforementioned mixture of multiple polypeptides of various lengths to be analyzed. The isotopically labeled standard may be added to the mixture as a complete sequence prior to the aforementioned protease digestion. Alternatively, a test sample of the mixture of polypeptides of various lengths and the isotopically labeled standard is digested with a protease in a separate reaction, and the protease-digested isotopically labeled standard is added to the test sample prior to analysis by LC / MS. The methods of the present invention further include quantifying the amount of the barcode fragment, the reference fragment, or both from the test sample by comparison to the quantification of the detected isotopically distinguishable amount of the barcode fragment, the reference fragment, or both.
[0057] Variations and modifications of these embodiments will occur to those skilled in the art after reviewing this disclosure. The features and aspects described above can be provided in any combination and subcombination (including multiple subcombinations and subcombinations) with one or more other features described herein. Various features described above or described above, including any components thereof, can be combined or integrated in other embodiments. Furthermore, certain features may be omitted or not provided.
[0058] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0059] The various features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure can be obtained by reference to the following detailed description and accompanying drawings that set forth illustrative embodiments, in which the principles of the invention are utilized. [Brief explanation of the drawings]
[0060] [Figure 1] Figure 1 shows a mixture of Xtenylated protease-activated T cell engager ("XPAT") polypeptides with XTEN polypeptides of various lengths. Full-length XPAT (top) contains a 288-amino acid XTEN polypeptide at the N-terminus and an 864-amino acid XTEN polypeptide at the C-terminus. Various truncations can occur in one or both of the N- and C-terminal XTEN polypeptides, for example, during fermentation, purification, or other steps in product preparation. Products with limited truncations (cleavage near a portion of the XTEN polypeptide distal to the protease-activated T cell engager linked to it) can function in a manner similar to full-length constructs, whereas severe truncations (cleavage near a portion of the XTEN polypeptide proximal to the protease-activated T cell engager linked to it) can have pharmacological properties significantly different from their full-length counterparts. The presence of truncations poses challenges in quantifying pharmacologically effective and ineffective variants in XPAT products. As shown in Figure 1 using full-length XPAT, each XTEN polypeptide has a proximal end and a distal end, where the proximal end is located closer to the biologically active polypeptide (e.g., a T cell engager, cytokine, monoclonal antibody (mAb), antibody fragment, or other protein being XTENed) relative to the distal end. Depending on the orientation of the linkage, the proximal or distal end of the XTEN polypeptide can correspond to the N-terminus or C-terminus of the XTEN polypeptide. [Figure 2]Figure 2 shows a mixture of XPAT polypeptides with barcoded XTEN polypeptides of various lengths. In full-length XPAT (top), the 288-amino acid-long N-terminal XTEN polypeptide contains three cleavably fused barcode sequences, "NA," "NB," and "NC," from the distal end to the proximal end, and the 864-amino acid-long C-terminal XTEN polypeptide contains three cleavably fused barcode sequences, "CC," "CB," and "CA," from the proximal end to the distal end. Each barcode is arranged to represent a pharmacologically relevant length of the corresponding XTEN polypeptide. For example, minor N-terminal truncation products of XPAT lacking the barcode "NA" but possessing the more proximal barcodes "NB" and "NC" can exhibit substantially the same pharmacological properties as the full-length construct. In contrast, the major N-terminal truncation products of XPAT, e.g., those lacking all three barcodes on the N-terminus, can have discernibly different pharmacological activity from the full-length construct. Unique proteolytic cleavable sequences are identified from biologically active XPAT polypeptides (here, tandem scFvs containing the active portion of a T cell engager). Because they are present in full-length variants of XPAT (including full-length XPAT, minor truncations, and major truncations), the unique proteolytic cleavable sequences can be used as references to quantify the amount of various cleavage products relative to the total amount of biologically active protein. [Figure 3]Figure 3 illustrates the potential design of barcoded XTEN polypeptides by inserting a barcode-generating sequence into a generic (or regular) XTEN polypeptide. An exemplary generic (or regular) XTEN polypeptide (top) contains non-overlapping 12-mer motifs within the sequence "BCDABDCDABDCBDCDABDCB" (where the sequence motifs "A," "B," "C," and "D" occur 3, 6, 5, and 7 times, respectively). A Glu-C protease digest of the exemplary generic XTEN polypeptide (top panel) does not yield unique peptides except at both ends ("NT" and "CT"). Insertion of a barcode-generating sequence, "X" (e.g., a unique 12-mer), into an XTEN polypeptide yields a unique proteolytically cleavable sequence (or barcode sequence) that does not occur anywhere else in the XTEN polypeptide. The barcode-generating sequence, "X," can be positioned so that the resulting barcode marks an XTEN polypeptide of pharmacologically relevant length. For example, an XTEN polypeptide that lacks a barcode can be functionally different from a corresponding XTEN polypeptide that has a barcode.Those skilled in the art will understand that the sequence that generates the barcode ("X") can be the barcode sequence itself.Alternatively, the sequence that generates the barcode ("X") can be different from the obtained barcode sequence.For example, the barcode sequence can overlap with, and therefore contain, a portion of the preceding or following 12-mer motif. [Figure 4A] 4A-4B illustrate quantification of N-terminal XTEN polypeptide cleavage levels. Figure 4A shows that barcoded XTEN polypeptides (bottom panel) can be constructed by replacing a sequence motif (e.g., D, the third sequence motif from the N-terminus) of a generic XTEN polypeptide (top panel) with a barcode-generating motif X; in this example, the barcode-generating motif ("X") itself is a unique proteolytically cleavable barcode sequence. As shown in the bottom panel of Figure 4A, the barcodes are positioned so that all severely cleaved forms of the XTEN polypeptide lack the barcode and all limitedly cleaved forms of the XTEN polypeptide contain the barcode. [Figure 4B] Figure 4B illustrates the relative abundance of various cleavage products in two different mixtures of XPAT. In one of the mixtures, the barcode is present in 99% of the constructs containing biologically active protein. In the other mixture, 13% of the constructs lack the barcode. Figures 4A-4B illustrate the use of barcoded XTEN polypeptides to distinguish between two polypeptide mixtures that have substantially similar average molecular weights but distinguishably different pharmacological activities. [Figure 5A] Figure 5A illustrates analytical size exclusion chromatography (SEC) of XPAT protein and the detection of the full-length protein and its cleaved derivatives. The synthetic protein and cleaved fractions contain fragments of the same size as the intact synthetic protein. [Figure 5B] Figure 5B illustrates the abundance of barcoded peptides in XPAT preparations as detected by mass spectrometry. Each measurement is the XIC range of the N-barcode SGPGSTPAE (SEQ ID NO: 8029) and C-barcode GSAPGTE (SEQ ID NO: 8023) normalized to a 400 nM spike of its corresponding heavy-isotope-labeled synthetic peptide.
[0061] The patent or application file contains at least one color drawing. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0062] term As used herein, the following terms have the meanings ascribed to them unless specified otherwise.
[0063] As used in this specification and claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "a cell" includes a plurality of cells, including mixtures thereof.
[0064] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. Polymers may be linear or branched, may contain modified amino acids, or may be interrupted by non-amino acids. These terms also encompass amino acid polymers modified by any other manipulation, such as disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, or conjugation with a labeling component.
[0065] As used herein, the term "amino acid" refers to any natural and / or unnatural or synthetic amino acid, including, but not limited to, glycine and both the D or L optical isomers, as well as amino acid analogs and peptidomimetics. Amino acids are designated using standard one-letter or three-letter codes.
[0066] A "host cell" includes an individual cell or cell culture that can be or has been a recipient of a human or animal vector. A host cell includes the progeny of a single host cell. The progeny will not necessarily be completely identical (in the form of the entire DNA complement or in the genome) to the original parent cell due to naturally occurring or genetically engineered variations.
[0067] A "chimeric" protein contains at least one polypeptide containing a region in a position within the sequence that is different from its naturally occurring position. The regions may normally be present in separate proteins and are brought together in the fusion polypeptide, or may normally be present in the same protein but are arranged in a new configuration in the fusion polypeptide. Such proteins may be described as "conjugated," "linked," "fused," or "fusion" proteins, which terms are used interchangeably herein and refer to the joining of two additional polypeptide sequences by any means, including chemical conjugation or recombinant means. Chimeric proteins can be made, for example, by chemical synthesis, or by creating and translating a polynucleotide in which the peptide regions are encoded in the desired relationship.
[0068] The terms "polynucleotide," "nucleic acid," "nucleotide," and "oligonucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides may have any three-dimensional structure and may perform any function known or to be discovered or developed. Polynucleotides may contain modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides may be further modified after polymerization, such as by conjugation with a labeling component.
[0069] The term "complementary strand of a polynucleotide" refers to a nucleotide molecule having a complementary base sequence and reverse orientation compared to a reference sequence, which is capable of hybridizing with complete fidelity to the reference sequence.
[0070] As used herein, polynucleotides that have "homology" or are "homologous" are those that hybridize under stringent conditions as defined herein and have at least 70%, preferably at least 80%, more preferably at least 90%, more preferably 95%, more preferably 97%, more preferably 98%, and even more preferably 99% sequence identity to these sequences.
[0071] The terms "percent identity" and "% identity," as applied to polynucleotide sequences, refer to the percentage of residue matches between at least two polynucleotide sequences aligned using a standardized algorithm. Such algorithms insert gaps in a standardized and reproducible manner within the sequences being compared to optimize the alignment between the two sequences, thereby achieving a more meaningful comparison of the two sequences. Percent identity can be measured over the length of the entire defined polynucleotide sequence, e.g., as defined by a particular SEQ ID NO:, or over a shorter length, e.g., over the length of a fragment taken from a longer defined polynucleotide sequence, e.g., at least 45, at least 60, at least 90, at least 120, at least 150, at least 210, or at least 450 contiguous residues. It is understood that such lengths are exemplary only, and that any fragment length supported by the sequences set forth in this specification, tables, figures, or sequence listing can be used to describe the length for which percent identity can be measured.
[0072] With respect to the polypeptide sequences identified herein, "percent (%) amino acid sequence identity" is defined as the percentage of amino acid residues in a query sequence that are identical to the amino acid residues of a second reference polypeptide sequence, or portion thereof, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity and without considering conservative substitutions as part of the sequence identity. Alignment to determine percent amino acid sequence identity can be achieved in a variety of ways within the skill of the art, for example, using publicly available computer software such as BLAST, BLAST-2, ALIGN, or Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms necessary to achieve maximum alignment across the entire length of the sequences being compared. Percent identity can be measured over the length of the complete defined polypeptide sequence, for example, as defined by a particular SEQ ID NO, or over a shorter length, for example, over the length of a fragment taken from a longer defined polypeptide sequence, for example, at least 15, at least 20, at least 30, at least 40, at least 50, at least 70, or at least 150 consecutive residues. Such lengths are exemplary only and may not be used in any manner whatsoever without reference to the present specification, tables, figures or sequences. It is understood that any fragment length supported by the sequences shown in the table can be used to describe the length over which the percentage identity can be determined.
[0073] As used herein, the "repetitiveness" of an XTEN polypeptide amino acid sequence refers to the repetitiveness of 3mers and can be measured by a computer program or algorithm or other means known in the art. The repetitiveness of 3mers of an XTEN polypeptide amino acid sequence can be assessed by determining the number of occurrences of overlapping 3mer sequences within the polypeptide. For example, a 200 amino acid residue polypeptide has 198 overlapping three-amino acid sequences (3mers), but the number of unique 3mer sequences depends on the amount of repetition within the sequence. A score (hereinafter referred to as a "subsequence score") can be generated that reflects the degree of 3mer repetitiveness within the overall polypeptide sequence. In the context of the present invention, "subsequence score" refers to the total occurrence of each unique 3mer frame across the 200 consecutive amino acid sequence of a polypeptide, divided by the absolute number of unique 3mer subsequences within the 200 amino acid sequence. Examples of such subsequence scores derived from the first 200 amino acids of repetitive and non-repetitive polypeptides are provided in Example 73 of International Patent Application Publication No. WO2010 / 091122A1, the entire contents of which are incorporated by reference. In some embodiments, the invention provides BPXTEN polypeptides, each comprising at least one XTEN polypeptide, wherein the XTEN polypeptide amino acid sequence can have a subsequence score of less than 16, or less than 14, or less than 12, or more preferably less than 10.
[0074] The term "substantially non-repetitive XTEN polypeptide amino acid sequence," as used herein, refers to an XTEN polypeptide in which there are few or no instances of four consecutive amino acids in the XTEN polypeptide amino acid sequence that are the same amino acid type, the XTEN polypeptide amino acid sequence has a subsequence score (as defined in the preceding paragraphs herein) of 12, or 10 or less, or there is no pattern in the N-terminal to C-terminal order of sequence motifs that make up the polypeptide sequence.
[0075] As used herein, the term "non-overlapping sequence motifs" includes completely non-overlapping sequence motifs as well as partially non-overlapping sequence motifs, with the proviso that the partially non-overlapping sequence motifs do not completely overlap.
[0076] A "vector" is a nucleic acid molecule, preferably one that is self-replicating in a suitable host, and transfers an inserted nucleic acid molecule into and / or between host cells. The term includes vectors that function primarily for the insertion of DNA or RNA into a cell, replicating vectors that function primarily for the replication of DNA or RNA, and expression vectors that function for the transcription and / or translation of DNA or RNA. Vectors that perform more than two of the above functions are also included. An "expression vector" is a polynucleotide that can be transcribed and translated into a polypeptide when introduced into a suitable host cell. An "expression system" usually connotes a suitable host cell containing an expression vector that can function to produce a desired expression product.
[0077] "t 1 / 2 " as used herein means ln(2) / K el The terminal half-life is calculated as K el is the terminal elimination rate constant calculated by linear regression of the terminal linear portion of the log concentration versus time curve. Half-life typically refers to the time required for half of the amount of an administered substance that accumulates in a living organism to be metabolized or eliminated by normal biological processes. "t 1 / 2 The terms "," "terminal half-life," "elimination half-life," and "circulating half-life" are used interchangeably herein.
[0078] The terms "antigen," "target antigen," or "immunogen" are used interchangeably herein. It refers to the structure or binding determinant to which an antibody fragment or antibody fragment-based therapeutic binds or has specificity for.
[0079] The term "payload" as used herein refers to a protein or peptide sequence having biological or therapeutic activity, a counterpart of a small molecule pharmacophore. Examples of payloads include, but are not limited to, cytokines, enzymes, hormones, and blood and growth factors. Payloads can also include genetically fused or chemically conjugated moieties such as chemotherapeutic agents, antiviral compounds, toxins, or imaging agents. These conjugated moieties can be attached to the remainder of the polypeptide via cleavable or non-cleavable linkers.
[0080] As used herein, "treatment" or "treating," "alleviating," and "alleviating" are used interchangeably herein and refer to an approach to obtaining a beneficial or desired result, including, but not limited to, a therapeutic benefit and / or a preventative benefit. "Therapeutic benefit" refers to the eradication or alleviation of the underlying disorder being treated. A therapeutic benefit is also achieved by the eradication or alleviation of one or more of the physiological symptoms associated with the underlying disease state, where improvement is observed in the human or animal, even though the human or animal may still be afflicted with the underlying disorder. For preventative benefit, the composition can be administered to a human or animal at risk of developing a particular disease state, or to a human or animal reporting one or more physiological symptoms of a disease, even if the disease has not been diagnosed.
[0081] "Therapeutic effect," as used herein, refers to a physiological effect caused by a fusion polypeptide of the invention, other than the ability to induce the production of antibodies against an antigenic epitope possessed by the biologically active protein, including, but not limited to, curing, alleviating, ameliorating, or preventing a disease state in a human or other animal, or to otherwise enhance the physical or mental well-being of a human or animal. Determination of a therapeutically effective amount is well within the capabilities of those skilled in the art, especially in light of the detailed description provided herein.
[0082] The terms "therapeutically effective amount" and "therapeutically effective dose," as used herein, refer to an amount of a biologically active protein, either alone or as part of a fusion protein composition, that, when administered in single or repeated doses to a human or animal, can produce some detectable beneficial effect on any symptom, aspect, measured parameter, or characteristic of a medical condition or disease state. Such an effect need not be absolute to be beneficial. A disease state can refer to a disorder or disease.
[0083] The term "therapeutically effective dose regimen," as used herein, refers to a schedule of continuously administered doses of a biologically active protein, either alone or as part of a fusion protein composition, wherein the doses are given in therapeutically effective amounts that provide a sustained benefit to any symptom, aspect, measured parameter, or characteristic of a pathology or disease state.
[0084] Fusion Polypeptides Disclosed herein are polypeptides comprising one or more extended recombinant polypeptides (XTEN or XTENs) (described more fully herein below) that may be fused or otherwise conjugated to another polypeptide, particularly a biologically active polypeptide; the foregoing embodiments are referred to herein as BPXTEN.
[0085] In some embodiments, the polypeptide is a first XTEN polypeptide ("extended XTEN polypeptide"). In some embodiments, the polypeptide further comprises a second XTEN polypeptide (such as those described below in the "Extended Recombinant Polypeptides (XTEN)" section or elsewhere herein). In some embodiments, the polypeptide comprises an XTEN polypeptide at or near its N-terminus (an "N-terminal XTEN"). In some embodiments, the polypeptide comprises an XTEN polypeptide at or near its C-terminus (a "C-terminal XTEN"). In some embodiments, the polypeptide comprises both an N-terminal XTEN polypeptide and a C-terminal XTEN polypeptide. In some embodiments, the first XTEN polypeptide is an N-terminal XTEN polypeptide and the second XTEN polypeptide is a C-terminal XTEN polypeptide.
[0086] The polypeptide can further comprise a biologically active polypeptide (“BP”) linked to the XTEN polypeptide, thereby forming an XTEN-containing fusion polypeptide referred to herein as a “BPXTEN” polypeptide.
[0087] The XTEN polypeptide can include one or more barcode fragments (described more fully below) that are releasable (formed to be released) from the XTEN polypeptide upon digestion of the fusion polypeptide (or BPXTEN) with a protease. In some embodiments, each barcode fragment differs in sequence and molecular weight from all other peptide fragments (including all other barcode fragments, if present) that are releasable from the polypeptide upon complete digestion of the polypeptide with a protease.
[0088] A (fusion) polypeptide can include one or more reference fragments (described more fully below) that are releasable (formed to be released) from the polypeptide upon protease digestion, e.g., to release barcode fragments from the polypeptide. In some embodiments, each reference fragment can be a single reference fragment that differs in sequence and molecular weight from all other peptide fragments that are releasable from the polypeptide upon digestion of the polypeptide with a protease.
[0089] Extended Recombinant Polypeptides (XTEN) Chain length and amino acid composition In some embodiments, the XTEN polypeptide comprises at least 150 amino acids. In some embodiments, the XTEN polypeptide is 150 to 3,000 amino acids in length, or 150 to 1,000 amino acids in length, or at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids in length. In some embodiments, at least 90% of the amino acid residues of the XTEN polypeptide are glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), or proline (P). In some embodiments, at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid residues of the XTEN polypeptide are selected from G, A, S, T, E, or P. In some embodiments, the XTEN polypeptide comprises at least four different types of G, A, S, T, E, or P amino acids. In some embodiments, the XTEN polypeptide comprises at least 150 amino acids, characterized in that at least 90% of the amino acid residues of the XTEN polypeptide are G, A, S, T, E, or P, and comprise at least four different types of amino acids selected from G, A, S, T, E, and P that are substantially randomized with respect to any other non-overlapping sequence motif that comprises the XTEN polypeptide. In some embodiments, the XTEN-containing fusion polypeptide (e.g., a fusion polypeptide comprising a biologically active polypeptide conjugated thereto) comprises a first XTEN polypeptide and a second XTEN polypeptide. In some embodiments, the sum of the total number of amino acids in the first XTEN polypeptide and the total number of amino acids in the second XTEN polypeptide is at least 300, at least 350, at least 400, at least 450, at least 50 ... At least 400, at least 500, at least 600, at least 700, or at least 800 amino acids.
[0090] Non-overlapping sequence motifs In some embodiments, the XTEN polypeptides provided herein comprise or are formed from multiple non-overlapping sequence motifs. In some embodiments, at least one of the non-overlapping sequence motifs is recurring (repeated at least twice within the XTEN) and at least one other of the non-overlapping sequence motifs is non-recurring (or occurs only once within the XTEN). In some embodiments, the multiple non-overlapping sequence motifs comprise a set of non-overlapping (non-recurring) sequence motifs, each of which is repeated at least twice within the XTEN and occurs (or does not occur) only once within the XTEN. In some embodiments, each non-overlapping sequence motif is 9-14 (or 10-14, or 11-13) amino acids in length. In some embodiments, each non-overlapping sequence motif is 12 amino acids in length. In some embodiments, the multiple non-overlapping sequence motifs comprise a set of non-overlapping (recurring) sequence motifs, each of which is repeated at least twice within the XTEN and is 9-14 amino acids in length. In some embodiments, the set of (recurring) non-overlapping sequence motifs comprises the 12-mer sequence motifs identified herein by SEQ ID NOs: 182-203 and 1715-1722 in Table 1. In some embodiments, the set of (recurring) non-overlapping sequence motifs comprises the 12-mer sequence motifs identified herein by SEQ ID NOs: 186-189 in Table 1. In some embodiments, the set of (recurring) non-overlapping sequence motifs comprises at least two, at least three, or all four of the 12-mer sequence motifs of SEQ ID NOs: 186-189 in Table 1. [Table 1]
[0091] Barcode fragment In some embodiments, the polypeptides provided herein comprise a barcode fragment (e.g., a first, second, or third barcode fragment of an XTEN polypeptide) that is releasable from the polypeptide upon digestion with a protease. In some embodiments, the barcode fragment comprises at least a portion of a sequence motif that occurs (or is found) only once within the XTEN (non-recurring, non-overlapping) and is distinct in sequence and molecular weight from all other peptide fragments that are releasable from the polypeptide upon complete digestion of the polypeptide with a protease. Those skilled in the art will understand that the term "barcode fragment" (or "barcode," or "barcode sequence") can refer to either the portion of the XTEN identified herein that is cleavably fused within the polypeptide or the resulting peptide fragment released from the polypeptide.
[0092] In some embodiments, the barcode fragment does not include the N-terminal or C-terminal amino acid of the XTEN polypeptide. As described in more detail below or elsewhere herein, in some embodiments, the barcode fragment is releasable (formed to be released) upon Glu-C digestion of the fusion polypeptide. In some embodiments, the barcode fragment does not include a glutamic acid immediately adjacent to another glutamic acid in the XTEN polypeptide. In some embodiments, the barcode fragment has a glutamic acid at its C-terminus. One of skill in the art will understand that the C-terminus of a barcode fragment, when cleavably fused within an XTEN polypeptide, can refer to the “final” (or most C-terminal) amino acid residue within the barcode fragment, even if other “non-barcode” amino acid residues are located C-terminal to the barcode fragment within the same XTEN polypeptide. In some embodiments, the barcode fragment has an N-terminal amino acid that immediately precedes a glutamic acid residue. In some embodiments, the glutamic acid residue preceding the N-terminal amino acid is not immediately adjacent to another glutamic acid residue. In some embodiments, the barcode fragment does not contain a glutamic acid residue at any position other than the C-terminus of the barcode fragment, unless the glutamic acid is immediately followed by a proline. In some embodiments, the barcode fragment is located 10-150, or 10-125 amino acids from either the N-terminus of the polypeptide or the C-terminus of the polypeptide. In some embodiments, the barcode fragment is located within 300, 280, 260, 250, 240, 220, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 48, 40, 36, 30, 24, 20, 12, or 10 amino acids or positions from the N-terminus of the polypeptide, or at any range between any of the foregoing. In some embodiments, the barcode fragment is located within 200, 150, 100, or 50 amino acids of the N-terminus of the polypeptide. In some embodiments, the barcode fragment is located between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the N-terminus of the polypeptide.In some embodiments, the barcode fragment is located within 300, 280, 260, 250, 240, 220, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 48, 40, 36, 30, 24, 20, 12, or 10 amino acids from the C-terminus of the polypeptide, or within any range between the foregoing. In some embodiments, the barcode fragment is located within 200, 150, 100, or 50 amino acids from the C-terminus of the polypeptide. In some embodiments, the barcode fragment is located between 10-200, 30-200, 40-150, or 50-100 amino acids from the C-terminus of the polypeptide. In some embodiments, the barcode fragment does not include the N-terminal amino acid or the C-terminal amino acid of the polypeptide, does not include a glutamic acid immediately adjacent to another glutamic acid in the XTEN, has a glutamic acid at its C-terminus, has an N-terminal amino acid immediately preceding the glutamic acid residue, and (v) is located 10-150, or 10-125 amino acids from either the N-terminus of the polypeptide or the C-terminus of the polypeptide. In some embodiments, the glutamic acid residue preceding the N-terminal amino acid is not immediately adjacent to another glutamic acid residue. In some embodiments, the barcode fragment does not include a glutamic acid residue at any position other than the C-terminus of the barcode fragment, unless the glutamic acid is immediately followed by a proline. In some embodiments, for a barcoded XTEN polypeptide fused to a biologically active polypeptide, at least one barcode fragment (or at least two barcode fragments, or three barcode fragments) contained in the barcoded XTEN is located at least 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, or 300 amino acids from the biologically active polypeptide. In some embodiments, the barcode fragment is at least 4, at least 5, at least 6, at least 7, or at least 8 amino acids in length. In some embodiments, the barcode fragment is at least 4 amino acids in length.In some embodiments, the barcode fragment is 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids in length, or within a range between any of the foregoing values. In some embodiments, the barcode fragment is selected from SEQ ID NOs: 8020-8030 (BAR001-BAR011) in Table 2. [Table 2]
[0093] In some embodiments, the barcoded XTEN polypeptide comprises only one barcode fragment. In some embodiments, the barcoded XTEN polypeptide comprises a set of barcode fragments, including a first barcode fragment, such as those described above or elsewhere herein. In these embodiments, each member of the set of barcode sequences is distinguishable from all other barcode sequences based on amino acid sequence or molecular weight (these methods of distinguishing between different barcode sequences will be relevant). In some embodiments, the set of barcode fragments comprises a second barcode fragment (or additional barcode fragments), such as those described above or elsewhere herein. In some embodiments, the set of barcode fragments comprises a third barcode fragment, such as those described above or elsewhere herein. The set of barcode fragments fused within an N-terminal XTEN polypeptide can be referred to as the N-terminal set of barcodes (the "N-terminal set"). The set of barcode fragments fused within a C-terminal XTEN polypeptide can be referred to as the C-terminal set of barcodes (the "C-terminal set"). In some embodiments, the N-terminal set comprises a first barcode fragment and a second barcode fragment. In some embodiments, the N-terminal set further comprises a third barcode fragment. In some embodiments, the C-terminal set comprises a first barcode fragment and a second barcode fragment. In some embodiments, the C-terminal set further comprises a third barcode fragment. In some embodiments, the second barcode fragment is located N-terminal to the first barcode fragment of the same set. In some embodiments, the second barcode fragment is located C-terminal to the first barcode fragment of the same set. In some embodiments, the third barcode fragment is located N-terminal to both the first and second barcode fragments. In some embodiments, the third barcode fragment is located C-terminal to both the first and second barcode fragments. In some embodiments, the third barcode fragment is located between the first and second barcode fragments.In some embodiments, the polypeptide comprises a set of barcode fragments comprising a first barcode fragment, a further (second) barcode fragment, and at least one additional barcode fragment, wherein each barcode fragment of the set of barcode fragments is part of a second XTEN polypeptide and differs in sequence and molecular weight from all other peptide fragments that are liberable from the polypeptide upon complete digestion of the polypeptide by a protease.
[0094] Exemplary Barcoded XTEN The amino acid sequences of 13 exemplary barcoded XTEN polypeptides containing one barcode (e.g., SEQ ID NOs: 8002-8003, 8005-8009, and 8013), two barcodes (e.g., SEQ ID NOs: 8001, 8004, 8010, and 8012), or three barcodes (e.g., SEQ ID NO: 8011) are set forth in Table 3a. Of these 13 exemplary barcoded XTEN polypeptides, six (SEQ ID NOs: 8001-8003, 8008-8009, and 8011) can be fused to a biologically active protein at the C-terminus of the biologically active protein, and seven (SEQ ID NOs: 8004-8007, 8010, and 8012-8013) can be fused to the N-terminus of a biologically active protein. In some embodiments, the XTEN polypeptide has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to a sequence selected from SEQ ID NOs: 8001-8019 of Table 3a. [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4] [Table 3-5]
[0095] In some embodiments, the barcoded XTEN polypeptides meet the following criteria: minimize sequence changes within the XTEN polypeptide; minimize changes in amino acid composition within the XTEN polypeptide; substantially maintain net changes in the XTEN polypeptide; Substantially maintaining (or improving) the low immunogenicity of the XTEN polypeptide and substantially maintaining (or improving) the pharmacokinetic properties of the XTEN polypeptide can be obtained by making one or more mutations in a generic XTEN polypeptide, such as any listed in Table 3b. In some embodiments, the amino acid sequence of the XTEN polypeptide has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 676-734 listed in Table 3b. In some embodiments, an XTEN sequence having at least 90% (e.g., at least 92%, at least 95%, at least 98%, or at least 99%) but less than 100% sequence identity to any one of SEQ ID NOs: 676-734 listed in Table 3b is obtained by making one or more mutations (e.g., fewer than 10, fewer than 8, fewer than 6, fewer than 5, fewer than 4, fewer than 3, or fewer than 2 mutations) to the corresponding sequence in Table 3b. In some embodiments, the one or more mutations comprise a deletion of a glutamic acid residue, an insertion of a glutamic acid residue, a substitution of a glutamic acid residue, or a substitution with a glutamic acid residue, or any combination thereof. In some embodiments, if the amino acid sequence of an XTEN polypeptide differs from any of SEQ ID NOs:676-734 listed in Table 3b but has at least 90% (e.g., at least 92%, at least 95%, at least 98%, or at least 99%) sequence identity, then at least 80%, at least 90%, at least 95%, at least 97%, or about 100% of the differences between the amino acid sequence of the XTEN polypeptide and the corresponding sequence in Table 3b comprise a deletion of a glutamic acid residue, an insertion of a glutamic acid residue, a substitution of a glutamic acid residue, or a substitution with a glutamic acid residue, or any combination thereof. In some such embodiments, at least 80%, at least 90%, at least 95%, at least 97%, or about 100% of the differences between the amino acid sequence of the XTEN polypeptide and the corresponding sequence in Table 3b include substitutions of glutamic acid residues, or substitutions for glutamic acid residues, or both.The term "first amino acid substitution," as used herein, refers to the substitution of a first amino acid residue with a second amino acid residue, resulting in the second amino acid residue occurring at the substitution position in the resulting sequence. For example, a "glutamic acid substitution" refers to the substitution of a glutamic acid (E) residue with a non-glutamic acid residue (e.g., serine (S)). The term "first amino acid substitution," as used herein, refers to the substitution of a second amino acid residue with a first amino acid residue, resulting in the first amino acid residue occurring at the substitution position in the resulting sequence. For example, a "glutamic acid substitution" refers to the substitution of a non-glutamic acid residue (e.g., serine (S)) with a glutamic acid residue. [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] [Table 4-7] [Table 4-8] [Table 4-9]
[0096] In some embodiments, to construct the sequence of a barcoded XTEN polypeptide, amino acid mutations are made on intermediate length XTEN polypeptides, to XTEN polypeptides of lengths greater than those of Table 3b, such as those in Table 3b, as well as those in which one or more 12-mer motifs of Table 1 have been added to the N- or C-terminus of the generic XTEN of Table 3b.
[0097] Additional examples of amino acid sequences of generic XTEN polypeptides that can be used in accordance with the present disclosure are described in U.S. Patent Publication Nos. 2010 / 0239554A1, 2010 / 0323956A1, 2011 / 0046060A1, 2011 / 0046061A1, and 2011 / 0077199A1, the disclosures of each of which are expressly incorporated herein by reference. or as described in International Patent Publication No. 2011 / 0172146A1, or International Patent Publication No. 2010091122A1, No. 2010144502A2, No. 2010144508A1, No. 2011028228A1, No. 2011028229A1, No. 2011028344A2, No. 2014 / 011819A2, or No. 2015 / 023891.
[0098] In some embodiments, a barcoded XTEN polypeptide fused within a polypeptide chain adjacent to the N-terminus of the polypeptide chain (an "N-terminal XTEN") can be attached to a His tag containing multiple poly(His) residues, including 6-8 His residues at the N-terminus, to facilitate purification of the fused polypeptide. In some embodiments, a barcoded XTEN polypeptide fused within a polypeptide chain at the C-terminus of the polypeptide chain (a "C-terminal XTEN polypeptide") can be included in or attached to the sequence EPEA at the C-terminus to facilitate purification of the fused polypeptide. In some embodiments, the fusion polypeptide comprises both an N-terminally barcoded XTEN polypeptide and a C-terminally barcoded XTEN polypeptide, wherein the N-terminally barcoded XTEN is linked to a His tag comprising multiple poly(His) residues, including 6-8 His residues at the N-terminus, and the C-terminally barcoded XTEN polypeptide is linked at the C-terminus to the sequence EPEA, thereby facilitating purification of the fusion polypeptide to a purity of, for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% by chromatographic methods known in the art, including, but not limited to, IMAC chromatography, C-tag XL affinity substrate, and other such methods, including, but not limited to, those described in the Examples section below.
[0099] Protease digestion The barcode fragments described above or elsewhere herein can be cleavably fused within an XTEN polypeptide and releasable (formed to be released) from the XTEN polypeptide upon digestion of the polypeptide by a protease. In some embodiments, the protease is a Glu-C protease. In some embodiments, the protease cleaves C-terminal to glutamic acid residues that are not followed by a proline. One skilled in the art will understand that barcoded XTEN polypeptides (XTEN polypeptides containing barcode fragments therein) are designed to achieve high efficiency, precision, and accuracy of protease digestion. For example, one skilled in the art will understand that adjacent Glu-Glu (EE) residues in an XTEN sequence can result in various cleavage patterns upon Glu-C digestion. Thus, when Glu-C protease is used for barcode release, the barcoded XTEN polypeptide or barcode fragment may not have a Glu-Glu (EE) sequence. Those skilled in the art will also understand that the di-peptide Glu-Pro (EP) sequence, when present in the fusion polypeptide, may not be cleavable by the Glu-C protease during the barcode release process.
[0100] Structural arrangement of BPXTEN In some embodiments, a BPXTEN fusion protein comprises a single BP polypeptide and a single XTEN polypeptide. Such a BPXTEN protein can have at least the following arrangements, listed from N- to C-terminus: BP-XTEN, XTEN-BP, BP-S-XTEN, and XTEN-S-BP (where "S" is a spacer sequence described below).
[0101] In some embodiments, the BPXTEN protein comprises a C-terminal XTEN polypeptide and, optionally, a spacer sequence (S) between the XTEN polypeptide and the BP polypeptide. Such a BPXTEN protein has Formula I (shown from N-terminus to C-terminus): (BP)-(S) x -(XTEN)(I), where BP is a biologically active protein as described herein below, S is a spacer sequence having from 1 to about 50 amino acid residues that may optionally include a BP release segment (described more fully herein below), x is either 0 or 1, and XTEN can be any XTEN polypeptide as described herein.
[0102] In some embodiments, the BPXTEN protein comprises an N-terminal XTEN polypeptide and, optionally, a spacer sequence (S) between the XTEN polypeptide and the BP protein. Such a BPXTEN protein has Formula II (shown from N-terminus to C-terminus): (XTEN)-(S) x -(BP)(II), where BP is a biologically active protein as described herein below, S is a spacer sequence having from 1 to about 50 amino acid residues that may optionally include a BP release segment (described more fully herein below), x is either 0 or 1, and XTEN can be any XTEN polypeptide as described herein.
[0103] In some embodiments, the BPXTEN protein comprises both an N-terminal XTEN polypeptide and a C-terminal XTEN polypeptide. Such a BPXTEN protein (e.g., XPAT in Figures 1-2) has the structure of Formula III: (XTEN)-(S) y -(BP)-(S) z -(XTEN)(III) where BP is a biologically active protein as described herein below and S is , a spacer sequence having 1 to about 50 amino acid residues that may optionally include a BP release segment (described more fully herein below), where y is either 0 or 1, z is either 0 or 1, and XTEN can be any XTEN polypeptide described herein.
[0104] Biologically active polypeptides Biologically active proteins (BPs) that can be fused to one or more XTEN polypeptides (described herein), particularly those disclosed herein below, including the sequences identified herein by Tables 4a-4h and 6a-6f, along with their corresponding nucleic acid and amino acid sequences, are well known in the art. Descriptions and sequences of these BPs are available in public databases, such as subscription-based databases such as Chemical Abstracts Services Databases (e.g., CAS Registry), GenBank, The Universal Protein Source (UniProt), and GenSeq (e.g., Derwent). A polynucleotide sequence encoding a BP can be a wild-type polynucleotide sequence that encodes a naturally occurring BP (e.g., either full-length or mature), or in some cases, the sequence can be a variant of the wild-type polynucleotide sequence (e.g., a polynucleotide that encodes a wild-type biologically active protein), e.g., the nucleotide sequence of the polynucleotide has been optimized for expression in a particular species, or can encode a variant of the wild-type protein, such as a site-directed mutation or an allelic variant. It is well within the capabilities of one of ordinary skill in the art to use wild-type or consensus cDNA sequences or codon-optimized variants of BP to generate BPXTEN constructs contemplated by the present invention using methods known in the art and / or in conjunction with the guidance and methods provided herein.
[0105] BPs for inclusion in the BPXTEN proteins disclosed herein (e.g., fusion polypeptides comprising at least one BP and at least one XTEN polypeptide) can include any protein of biological, therapeutic, prophylactic, or diagnostic benefit or function, or that, when administered to a human or animal, are useful for maintaining biological activity or for preventing or alleviating a disease, disorder, or condition. Particularly advantageous are BPs that seek to enhance pharmacokinetic parameters, increase solubility, increase stability, mask activity, or some other pharmaceutical property, or that have a longer terminal half-life to improve efficacy, safety, or result in less frequent dosing and / or improve patient compliance, compared to a BP not linked to an XTEN polypeptide. Thus, BPXTEN fusion protein compositions can be prepared with a variety of objectives in mind, including improving the therapeutic efficacy of a biologically active compound when administered to a human or animal, compared to a BP not linked to an XTEN polypeptide, for example, by increasing the in vivo exposure or the length of time that BPXTEN remains in the therapeutic window.
[0106] The BP can be a naturally occurring full-length protein, or can be a biologically active fragment or sequence variant of the protein that retains at least some of the biological activity of the naturally occurring protein.
[0107] In one embodiment, the BP incorporated into a human or animal composition can be a recombinant polypeptide having a sequence corresponding to a protein found in nature. In another embodiment, the BP can be a sequence variant, fragment, homolog, or mimetic of a naturally occurring sequence that retains at least some of the biological activity of the naturally occurring BP. In a non-limiting example, the BP can have at least about 80% sequence identity, or alternatively, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, to a protein sequence selected from Tables 4a-4h. can be sequences exhibiting 100% sequence identity. In a further non-limiting example, a BP can be a bispecific sequence comprising a first binding domain and a second binding domain, wherein the first binding domain having specific binding affinity for a tumor-specific marker or antigen of a target cell has at least about 80% sequence identity, or alternatively, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96% or more of a sequence identity with the paired VL and VH sequences of an anti-CD3 antibody identified in Table 6f. , 97%, 98%, 99%, or 100% sequence identity and the second binding domain having specific binding affinity for effector cells exhibits at least about 80% sequence identity, or alternatively, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, to the paired VL and VH sequences of the anti-target cell antibodies identified in Table 6a. In one embodiment, a BPXTEN fusion protein can comprise a single BP protein linked to an XTEN polypeptide. In another embodiment, a BPXTEN protein can comprise a first BP and a second molecule of the same BP, resulting in a fusion protein comprising two BPs (e.g., two molecules of glucagon or two molecules of hGH) linked to one or more XTEN polypeptides.
[0108] Generally, when used in vivo or employed in an in vitro assay, a BP exhibits binding specificity or another desired biological characteristic for a predetermined target (or a predetermined number of targets). For example, the BP can be an agonist, receptor, ligand, antagonist, enzyme, antibody (e.g., mono- or bispecific), or hormone. Of particular interest are BPs used or known to be useful for diseases or disorders in which the native BP has a relatively short terminal half-life and enhanced pharmacokinetic parameters (which can optionally be released from the fusion protein by cleavage of a spacer sequence) allow for less frequent dosing or enhanced pharmacological effect. Also of interest are BPs that have a narrow therapeutic window between the minimum effective dose or blood concentration (Cmin) and the maximum tolerated dose or blood concentration (Cmax). In such cases, linking the BP to a fusion protein containing a selected XTEN polypeptide sequence can result in improved properties, making it more useful as a therapeutic or prophylactic agent compared to a BP that is not linked to one or more XTEN polypeptides.
[0109] Glucose-regulating peptides Endocrine and obesity-related diseases or disorders have reached epidemic proportions in most developed countries and represent a substantial and increasing healthcare burden in most developed countries, including a wide variety of conditions affecting the body's organs, tissues, and circulatory system. Of particular concern are endocrine and obesity-related diseases and disorders, among them diabetes, which is one of the leading causes of death in the United States.
[0110] Most metabolic processes in glucose homeostasis and insulin response are regulated by multiple peptides and hormones, and many such peptides and hormones, as well as their analogs, have found utility in the treatment of metabolic diseases and disorders. Many of these peptides tend to be highly homologous to one another, even when they have opposing biological functions. Glucose-increasing peptides are exemplified by the peptide hormone glucagon, while glucose-lowering peptides include exendin-4, glucagon-like peptide 1, and amylin. However, the use of therapeutic peptides and / or hormones, even when augmented by the use of small molecule drugs, has met with limited success in managing such diseases and disorders. In particular, dose optimization is important for drugs and biologics used in the treatment of metabolic diseases, especially those with a narrow therapeutic window. Hormones in general, and peptides involved in glucose homeostasis, often have a narrow therapeutic window. A narrow therapeutic window is often due to the fact that such hormones and peptides, Combined with the fact that therapeutic proteins typically have short half-lives, frequent dosing is required to achieve clinical benefit, making the management of such patients difficult. While chemical modifications to therapeutic proteins, such as pegylation, can modify their in vivo clearance rate and subsequent serum half-life, they require additional manufacturing steps and result in heterogeneous final products. In addition, unacceptable side effects from chronic administration have been reported. Alternatively, genetic modification by fusion of an Fc domain to a therapeutic protein or peptide increases the size of the therapeutic protein, reducing its clearance rate through the kidney and promoting its recycling from lysosomes via the FcRn receptor. Unfortunately, Fc domains tend to fold inefficiently during recombinant expression and form insoluble precipitates known as inclusion bodies. These inclusion bodies must be solubilized, and functional proteins must be refolded, a time-consuming, inefficient, and expensive process.
[0111] Thus, one aspect of the present invention is the incorporation of peptides involved in glucose homeostasis, insulin resistance, and obesity (collectively, "glucose-regulating peptides") to create compositions with utility in the treatment of glucose, insulin, and obesity disorders, diseases, and related conditions. Suitable glucose-regulating peptides that can be linked to the XTEN polypeptides disclosed herein to create BPXTEN proteins include all biologically active polypeptides, particularly peptides that increase glucose-dependent secretion of insulin by pancreatic beta cells or enhance insulin action. Glucose-regulating peptides can also include biologically active polypeptides that stimulate pro-insulin gene transcription in pancreatic beta cells. Furthermore, glucose-regulating peptides can also include biologically active polypeptides that slow gastric emptying and reduce food intake. Glucose-regulating peptides can also include biologically active polypeptides that inhibit glucagon release from the alpha cells of the islets of Langerhans. Table 4a provides a non-limiting list of glucose-regulating peptide sequences that can be encompassed by the BPXTEN fusion proteins of the present invention. The glucose-regulating peptides of the BPXTEN compositions disclosed herein can be peptides that exhibit at least about 80% sequence identity (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity) to an amino acid sequence selected from Table 4a. [Table 5-1] [Table 5-2]
[0112] "Adrenomedullin" or "ADM" refers to the human adrenomedullin peptide hormone and species and sequence variants thereof that possess at least some of the biological activity of mature ADM. ADM is derived from a 185 amino acid preprohormone through sequential enzymatic cleavage and amidation, resulting in a 52 amino acid prohormone with a measured plasma half-life of 22 minutes. This results in a biologically active peptide of the formula: ADM. The ADM-containing fusion proteins of the present invention may find particular use in diabetes for glucose regulation or for their stimulatory effect on insulin secretion from pancreatic islet cells in humans or animals with persistent hypotension. The complete genomic basis of human AM has been reported (Ishimitsu et al., 1994, Biochem. Biophys. Res. Commun 203:631-639), and analogs of the ADM peptide have been cloned as described in U.S. Patent No. 6,320,022.
[0113] "Amylin" refers to a human peptide hormone called amylin, pramlintide, and its species variations described in U.S. Patent No. 5,234,906, which possess at least some of the biological activity of mature amylin. Amylin is a 37-amino acid polypeptide hormone co-secreted with insulin by pancreatic beta cells in response to nutrient intake (Koda et al., 1992, Lancet 339:1179-1180) and has been reported to regulate several key pathways of carbohydrate metabolism, including glucose incorporation into glycogen. Amylin-containing fusion proteins of the present invention can complement the action of insulin and regulate the rate of glucose clearance from the circulation and its uptake by peripheral tissues. Amylin analogs are cloned as described in U.S. Patent Nos. 5,686,411 and 7,271,238.
[0114] Amylin mimetics can be made that retain biological activity. For example, pramlintide has the sequence KCNTATCATNRLANFLVHSSNNFGPILPPTNVGSNTY (SEQ ID NO: 43), where amino acids from the rat amylin sequence are identical to those in the human amylin sequence. In one embodiment, the present invention provides a polypeptide having the sequence KCNTATCATX1RLANFLVHSSNNFGX2ILX2X2TNVGSNTY (SEQ ID NO: 44), wherein X1 is independently N or Q, and X2 is independently S, P, or G. In one embodiment, the amylin mimetic incorporated into BPXTEN may have the sequence KCNTATCATNRLANFLVHSSNNFGGILGGTNVGSNTY (SEQ ID NO: 45). In another embodiment, where the amylin mimetic is used at the C-terminus of BPXTEN, the mimetic may have the sequence KCNTATCATNRLANFLVHSSNNFGGILGGTNVGSNTY(NH2) (SEQ ID NO: 46).
[0115] "Calcitonin" (CT) refers to the human calcitonin protein and its species and sequence variants, including salmon calcitonin ("sCT"), which possess at least some of the biological activity of mature CT. CT is a 32-amino acid peptide cleaved from a larger thyroid prohormone that is thought to function in the nervous and vascular systems but has also been reported to be a potent hormonal mediator of the satiety reflex. (Reviewed in Becker, JCEM, 89(4):1512-1525 (2004) and Sexton, Current Medicinal Chemistry 6:1067-1093 (1999)). Calcitonin-containing fusion proteins of the invention may find particular use for the treatment of osteoporosis and as a therapy for Paget's disease of bone. Synthetic calcitonin peptides are produced as described in U.S. Pat. Nos. 5,175,146 and 5,364,840.
[0116] "Calcitonin gene-related peptide" or "CGRP" refers to the human CGRP peptide, a member of the calcitonin family of peptides, which exists in humans in two forms: α-CGRP (a 37-amino acid peptide) and β-CGRP. It also refers to species and sequence variants thereof that possess at least some of the biological activity of mature CGRP. CGRP shares 43-46% sequence identity with human amylin. CGRP-containing fusion proteins of the invention may find particular use in reducing morbidity associated with diabetes, alleviating hyperglycemia and insulin deficiency, inhibiting lymphocyte infiltration into pancreatic islets, and protecting beta cells against autoimmune destruction. Methods for producing synthetic and recombinant CGRP are described in U.S. Patent No. 5,374,618.
[0117] "Cholecystokinin" or "CCK" refers to the human CCK peptide and its species and sequence variants that have at least some of the biological activity of mature CCK. CCK-58 is the mature sequence, while the CCK-33 amino acid sequence first identified in humans is the major circulating form of the peptide. The CCK family also includes the eight amino acid in vivo C-terminal fragment ("CCK-8"), CCK-5, the pentagastrin or C-terminal peptide CCK(29-33), and CCK-4, the C-terminal tetrapeptide CCK(30-33). CCK is a gastrointestinal peptide hormone involved in stimulating the digestion of fats and proteins. The CCK-33 and CCK-8-containing fusion proteins of the present invention may find particular use in reducing the increase in circulating glucose and enhancing the increase in circulating insulin after meal ingestion. Analogs of CCK-8 are prepared as described in U.S. Patent No. 5,631,230.
[0118] "Exendin-3" refers to a glucose-regulating peptide isolated from the Mexican beaded lizard and sequence variants thereof that have at least some of the biological activity of mature exendin-3. Exendin-3 amide is a specific exendin receptor antagonist derived from it that mediates an increase in pancreatic cAMP and the release of insulin and amylase. Exendin-3-containing fusion proteins of the present invention are useful for the treatment of diabetes and insulin resistance. The assay may find particular use in the treatment of insulin-resistant disorders. Sequences and methods for the assay are described in U.S. Patent No. 5,424,286.
[0119] "Exendin-4" refers to a glucose-regulating peptide found in the saliva of the Gila monster Gila monster, as well as species and sequence variants thereof, including the naturally occurring 39-amino acid sequence HGEGTFTSDLSKQMEEEAVRLFIEYLKNGGPSSGAPPPS (SEQ ID NO: 47) and its homologous sequences and peptidomimetics, and variants, naturally occurring and non-naturally occurring sequences, such as those derived from primates, that have at least some of the biological activity of mature exendin-4. Exendin-4 is an incretin polypeptide hormone that increases blood glucose, stimulates insulin secretion, delays gastric emptying, and improves satiety, thereby resulting in significant improvement in postprandial hyperglycemia. Table 4b presents sequences from a variety of species, while Table 4c presents a list of synthetic GLP-1 analogs, all of which are contemplated for use in the BPXTEN proteins described herein.
[0120] Fibroblast growth factor 21, or "FGF-21," refers to the human protein encoded by the FGF-21 gene, or species and sequence variants thereof that have at least some of the biological activity of mature FGF-21. FGF-21 stimulates glucose uptake in adipocytes but not other cell types, an effect that is additive to insulin activity. FGF-21-containing fusion proteins of the present invention may find particular use in the treatment of diabetes, including by inducing energy expenditure, fat utilization, and lipid excretion. FGF-21 is cloned as described in U.S. Patent No. 6,716,626.
[0121] Fibroblast growth factor 19, or "FGF-19," refers to the human protein encoded by the FGF-19 gene, or species and sequence variants thereof that have at least some of the biological activity of mature FGF-19. FGF-19 is a member of the fibroblast growth factor (FGF) family of proteins. It increases hepatic expression of the leptin receptor, metabolic rate, stimulates glucose uptake in adipocytes, and leads to weight loss in obese mouse models (Fu et al., 2004, Endocrinology 145:2504-2603). FGF-19-containing fusion proteins of the present invention may find particular use in increasing metabolic rate and reversing diet-induced and leptin-deficient diabetes. FGF-19 was cloned and expressed as described in U.S. Patent Publication No. 20020042367.
[0122] "Gastrin" refers to human gastrin peptides, truncated versions, and species and sequence variants that have at least some of the biological activity of mature gastrin. Gastrin is found primarily in three forms: gastrin-34 ("large gastrin"), gastrin-17 ("small gastrin"), and gastrin-14 ("minimal gastrin"), which share sequence homology with CCK. The gastrin-containing fusion proteins of the present invention may find particular use in the treatment of obesity and diabetes for glucose regulation. Gastrin is synthesized as described in U.S. Patent No. 5,843,446.
[0123] "Ghrelin" refers to the human hormone that induces satiety, or species and sequence variants, including the naturally occurring processed 27 or 28 amino acid sequence and homologous sequences. Ghrelin levels increase before meals and decrease after meals, and may increase food intake and fat mass through actions exerted at the level of the hypothalamus. The ghrelin-containing fusion proteins of the present invention may find particular use as agonists, for example, to selectively stimulate GI tract motility in gastrointestinal motility disorders, to promote gastric emptying, or to stimulate growth hormone release. Sequences such as those described in U.S. Pat. No. 7,385,026 may be used. Ghrelin analogs or truncated variants with substitutions may find particular use as fusion partners with XTEN polypeptides for use as antagonists of improved glucose homeostasis, for the treatment of insulin resistance, and for the treatment of obesity. The isolation and characterization of ghrelin have been reported (Kojima et al., 1999, Nature. 402:656-660), and synthetic analogs are prepared by peptide synthesis as described in U.S. Patent No. 6,967,237.
[0124] "Glucagon" refers to human glucagon glucose-regulating peptide, or species and sequence variants thereof, including the naturally occurring 29 amino acid sequence and homologous sequences, naturally occurring sequence variants such as those derived from primates, and non-naturally occurring sequence variants, which have at least a portion of the biological activity of mature glucagon. The term "glucagon" as used herein also includes peptidomimetics of glucagon. Glucagon-containing fusion proteins of the present invention may find particular use in increasing blood glucose levels in individuals with existing hepatic glycogen stores and maintaining glucose homeostasis in diabetes. Glucagon is cloned as described in U.S. Pat. No. 4,826,763.
[0125] "GLP-1" refers to human glucagon-like peptide-1 and its sequence variants that have at least some of the biological activity of mature GLP-1. The term "GLP-1" includes human GLP-1(1-37), GLP-1(7-37), and GLP-1(7-36)amide. GLP-1 stimulates insulin secretion, but only during periods of hyperglycemia. The safety of GLP-1 compared to insulin is enhanced by this property and the knowledge that the amount of insulin secreted is a fraction of the magnitude of hyperglycemia. The biological half-life of GLP-1(7-37)OH is only 3-5 minutes (U.S. Patent No. 5,118,666). GLP-1-containing fusion proteins of the present invention may find particular use in the treatment of diabetes and insulin resistance disorders for glucose regulation. GLP-1 may be cloned and derivatives prepared as described in U.S. Patent No. 5,118,666. Non-limiting examples of GLP-1 sequences from a wide variety of species are shown in Table 4b, while Table 4c shows the sequences of numerous synthetic GLP-1 analogs, all of which are contemplated for use in the BPXTEN compositions described herein. [Table 6] [Table 7-1] [Table 7-2] [Table 7-3]
[0126] The GLP native sequence can be described by several sequence motifs, presented below, where the letters in brackets represent the permissible amino acids at each sequence position: {HVY}{AGISTV}{DEHQ}{AG}{ILMPSTV}{FLY}{DINST}{ADEKNST}{ADENSTV}{LMVY}{ANRSTY}{EHIKNQRST}{AHILMQVY}{LMRT}{ADEGKQS}{ADEGKNQSY}{AEIKLMQR}{AKQRSVY}{{AILMQSTV}{GKQR}{DEKLQR}{FHLVWY}{ILV}{ADEGHIKNQRST}{ADEGNRSTW}{GILVW}{AIKLMQSV}{ADGIKNQRST}{GKRSY} (SEQ ID NO: 9399). In addition, synthetic analogs of GLP-1 may be useful as fusion partners to XTEN polypeptides to generate BPXTEN proteins with biological activity useful in the treatment of glucose-related disorders.
[0127] "GLP-2" refers to human glucagon-like peptide-2 and sequence variants thereof that have at least some of the biological activity of mature GLP-2. More specifically, GLP-2 is a 33 amino acid peptide that is co-secreted with GLP-1 from enteroendocrine-mediated cells in the small and large intestine.
[0128] "Insulin-like growth factor 1" or "IGF-1" refers to the human IGF-1 protein and species and sequence variants thereof that have at least some of the biological activity of mature IGF-1. IGF-1 consists of 70 amino acids and is primarily produced by the liver as an endocrine hormone and in target tissues in a paracrine / autocrine manner. IGF-1-containing fusion proteins of the present invention may find particular use in the treatment of diabetes and insulin resistance disorders for glucose regulation. IGF-1 has been cloned and expressed in E. coli and yeast as described in U.S. Pat. No. 5,324,639.
[0129] "Insulin-like growth factor 2" or "IGF-2" refers to the human IGF-2 protein and species and sequence variants thereof that have at least some of the biological activity of mature IGF-2. IGF-2 is a protein derived from the human IGF-2 protein, as described by Bell et al., 1985, Proc Natl Cloned as described in Acad Sci USA. 82:6450-4.
[0130] "Islet neogenesis associated protein" (INGAP), or "pancreatic beta cell growth factor," is the human INGAP peptide and species and sequence variants that possess at least some of the biological activity of mature INGAP. INGAP-containing fusion proteins of the invention may find particular use in the treatment or prevention of diabetes and insulin resistance disorders. INGAP was cloned and expressed as described by R Rafaeloff et al., 1997, J Clin Invest. 99(9):2100-2109.
[0131] "Intermedin" or "AFP-6" refers to the human intermedin peptide as well as species and sequence variants thereof that have at least some of the biological activity of mature intermedin. Intermedin refers to a peptide that inhibits gastric emptying and reduces blood pressure in both normal and hypertensive humans or animals, and is associated with glucose homeostasis. The intermedin-containing fusion proteins of the present invention may find particular use in the treatment of diabetes, insulin resistance disorders, and obesity. Intermedin peptides and variants are cloned as described in U.S. Pat. No. 6,965,013.
[0132] "Leptin" refers to naturally occurring leptin from any species, as well as biologically active D-isoforms, or fragments and sequence variants thereof. Leptin-containing fusion proteins of the present invention may find particular use in the treatment of diabetes, insulin resistance disorders, and obesity for glucose regulation. Leptin has been cloned as described in U.S. Pat. No. 7,112,659, and leptin analogs and fragments have been cloned as described in U.S. Pat. No. 5,521,283, U.S. Pat. No. 5,532,336, PCT / US96 / 22308, and PCT / US96 / 01471.
[0133] "Neuromedin" refers to the neuromedin family of peptides, including neuromedin U and S peptides, as well as sequence variants thereof. Various truncated or spliced variants, such as FLFHYSKTQKLGKSNVVEELQSPFASQSRGYFLFRPRN (SEQ ID NO: 180), are included in the neuromedin U family. An example of the neuromedin S family is human neuromedin S, particularly its amide form, having the sequence ILQRGSGTAAVDFTKKDHTATWGRPFFLFRPRN (SEQ ID NO: 181). The neuromedin fusion proteins of the present invention may find particular use in the treatment of obesity, diabetes, reducing food intake, and other related conditions and disorders described herein.
[0134] "Oxyntomodulin," or "OXM," refers to human oxyntomodulin as well as species and sequence variants that have at least some of the biological activity of mature OXM. OXM is a 37-amino acid peptide produced in the colon that contains the 29-amino acid sequence of glucagon followed by an 8-amino acid carboxy-terminal extension. OXM-containing fusion proteins of the invention may find particular use in the treatment of diabetes, insulin resistance, and obesity for glucose regulation, and can be used as a treatment for weight loss.
[0135] "PYY" refers to human peptide YY polypeptides, as well as species and sequence variants having at least some of the biological activity of mature PYY. PPY-containing fusion proteins of the invention may find particular use in the treatment of diabetes, insulin resistance disorders, and obesity for glucose regulation. PYY analogs are prepared as described in U.S. Patent Nos. 5,604,203, 5,574,010, and 7,166,575.
[0136] "Urocortin" refers to human urocortin peptide hormones and sequence variants thereof that have at least some of the biological activity of mature urocortin. Three human urocortins exist: Ucn-1, Ucn-2, and Ucn-3. Additional urocortins and analogs are described in U.S. Pat. No. 6,214,797. The urocortin-containing BPXTEN proteins of the present invention may also find particular use in the treatment or prevention of conditions associated with stimulated ACTH release, hypertension due to vasodilatory effects, inflammation mediated through pathways other than elevated ACTH, hyperthermia, appetite disorders, congestive heart failure, stress, anxiety, and psoriasis. Urocortin-containing fusion proteins can also be combined with natriuretic peptide modules, amylin family and exendin family modules, or GLP1 family modules to enhance cardiovascular benefits, e.g., treatment of CHF by providing beneficial vasodilatory effects.
[0137] Metabolic Disease and Cardiovascular Proteins Metabolic and cardiovascular diseases impose a substantial healthcare burden in most developed countries, and cardiovascular disease remains the leading cause of death and disability in the United States and most European countries. Metabolic diseases and disorders include a wide variety of conditions that affect the body's organs, tissues, and circulatory system.
[0138] Dyslipidemia occurs frequently in humans or animals with diabetes and cardiovascular disease, typically characterized by parameters such as elevated plasma triglycerides, low HDL (high-density lipoprotein) cholesterol, normal to elevated levels of LDL (low-density lipoprotein) cholesterol, and elevated levels of low-density LDL particles in the blood. Dyslipidemia and hypertension are major contributors to the increased incidence of coronary events, renal disease, and death in humans or animals with metabolic diseases such as diabetes and cardiovascular disease.
[0139] Cardiovascular disease may be manifested by a number of disorders, symptoms and clinical parameters involving the heart, vasculature and organ systems throughout the body, including aneurysms, angina, atherosclerosis, cerebrovascular accident (stroke), cerebrovascular disease, congestive heart failure, coronary artery disease, myocardial infarction, reduced cardiac output and peripheral vascular disease, hypertension, hypotension, blood markers (e.g., C-reactive protein, BNP, and enzymes such as CPK, LDH, SGPT, SGOT), among others.
[0140] Most metabolic processes and many cardiovascular parameters are regulated by multiple peptides and hormones ("metabolic proteins"), and many such peptides and hormones, as well as their analogs, have found utility in the treatment of such diseases and disorders. However, the use of therapeutic peptides and / or hormones, even when augmented by the use of small molecule drugs, has limited success in managing such diseases and disorders. In particular, dose optimization is important for drugs and biologics used in the treatment of metabolic diseases, especially those with a narrow therapeutic window. Hormones in general, and peptides involved in glucose homeostasis, often have a narrow therapeutic window. This narrow therapeutic window, coupled with the fact that such hormones and peptides typically have a short half-life, requires frequent dosing to achieve clinical benefit, making the management of such patients difficult. Therefore, there remains a need for treatments with increased efficacy and safety in the treatment of metabolic diseases.
[0141] Thus, one aspect of the present invention is the incorporation of biologically active metabolic proteins involved in or used in the treatment of metabolic and cardiovascular diseases and disorders into BPXTEN fusion proteins to create compositions useful in the treatment of such disorders, diseases, and related conditions. Metabolic proteins can include any protein with a biological, therapeutic, or prophylactic benefit or function that is useful in the prevention, treatment, intervention, or alleviation of metabolic or cardiovascular diseases, disorders, or conditions. Table 4d provides a non-limiting list of such sequences of metabolic BPs encompassed by the BPXTEN fusion proteins of the present invention. The metabolic protein of the BPXTEN compositions of the present invention can be a protein exhibiting at least about 80% sequence identity, or alternatively, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a protein sequence selected from Table 4d. [Table 8]
[0142] "Anti-CD3" refers to the monoclonal antibody against the T cell surface protein CD3, OKT3 (also called muromonab), and the humanized anti-CD3 monoclonal antibody (hOKT31(Ala-Ala)) (Herold et al., 2002, New England "Anti-CD3" refers to species and sequence variants, and fragments thereof, including those of the "BPXTEN" family of fusion proteins (Journal of Medicine 346:1692-1698). The anti-CD3-containing fusion proteins of the present invention may find particular use for delaying new-onset type 1 diabetes, including the use of anti-CD3 as a therapeutic effector and targeting moiety for a second therapeutic BP in BPXTEN compositions. The variable region sequences and the production of anti-CD3 are described in U.S. Pat. No. 5,885,525. ,573 and 6,491,916.
[0143] "IL-1ra" refers to species and sequence variants, including the human IL-1 receptor antagonist protein and the sequence variant anakinra (Kineret®), which possesses at least some of the biological activity of mature IL-1ra. Anakinra is a non-glycosylated recombinant human IL-1ra that differs from endogenous human IL-1ra by the addition of an N-terminal methionine. A commercialized version of anakinra is sold as Kineret®. It binds to the IL-1 receptor with the same avidity as native IL-1ra and IL-1b, but does not result in receptor activation (signal transduction), an effect attributed to the presence of only one receptor-binding motif on IL-1ra versus two such motifs on IL-1α and IL-1β. Anakinra has 153 amino acids, a size of 17.3 kD, and a reported half-life of approximately 4-6 hours.
[0144] Increased IL-1 production has been reported in patients with various microbial infectious diseases and various other diseases. The IL-1ra-containing fusion proteins of the present invention may find particular use in treating any of the aforementioned diseases and disorders. IL-1ra is cloned as described in U.S. Patent Nos. 5,075,222 and 6,858,409.
[0145] "Natriuretic peptide" refers to atrial natriuretic peptide (ANP), brain natriuretic peptide (BNP or B-type natriuretic peptide), and C-type natriuretic peptide (CNP), both human and non-human species and sequence variants thereof, which possess at least some of the biological activity of the mature counterpart natriuretic peptides. Sequences of useful forms of natriuretic peptides are disclosed in U.S. Patent Publication No. 20010027181. Examples of ANP include those from various species, including human ANP (Kangawa et al., 1984, BBRC 118:131) or porcine and rat ANP (Kangawa et al., 1984, BBRC 121:585). Sequence analysis revealed that preproBNP consists of 134 residues and is cleaved to the 108 amino acid proBNP. Cleavage of a 32 amino acid sequence from the C-terminus of ProBNP results in the circulating physiologically active form, human BNP(77-108). The 32 amino acid human BNP is involved in the formation of disulfide bonds (Sudoh et al., 1989, BBRC 159:1420) and U.S. Patent Nos. 5,114,923, 5,674,710, 5,674,710, and 5,948,761. BPXTEN, which has one or more natriuretic functions, may be useful in treating hypertension, inducing diuresis, inducing natriuresis, widening or relaxing vascular conduction, binding natriuretic peptide receptors (e.g., NPR-A), inhibiting secretion of aldosterone from the adrenal gland, treating cardiovascular diseases and disorders, arresting or reversing cardiac tissue repair after a cardiac event or as a result of congestive heart failure, treating renal diseases and disorders, treating or preventing ischemic stroke, and treating asthma.
[0146] "Heparin-binding growth factor 2" or "FGF-2" refers to the human FGF-2 protein, as well as species and sequence variants thereof that have at least some of the biological activity of the mature counterpart. FGF-2 is cloned as described in Burgess, W.H. and Maciag, T., Ann. Rev. Biochem., 58:575-606 (1989); Coulier, F., et al., 1994, Prog. Growth Factor Res. 5:1, and PCT Publication No. 87 / 01728.
[0147] "TNF receptor" refers to the human receptor for TNF, as well as species and sequence variants thereof that possess at least some of the biological receptor activity of the mature TNFR. The X-ray crystal structure of the complex formed by the NF receptor and the extracellular domain of TNFβ has been determined (Banner et al., 1993 Cell 73:431, incorporated herein by reference).
[0148] clotting factors In hemophilia, blood clotting is impaired by the absence of certain plasma blood clotting factors. Human factor IX (FIX) is a serine protease zymogen that is a key component of the intrinsic pathway of the blood clotting cascade. Factor VIIa (FVIIa) protein has found utility in treating bleeding episodes in patients with hemophilia A or B and patients with acquired hemophilia using inhibitors of FVIII or FIX, as well as for surgical or invasive procedures in patients with hemophilia A or B using inhibitors of FVIII or FIX. Thus, there remains a need for factor IX and factor VIIa compositions with extended half-life and maintained activity when administered as part of a prophylactic and / or therapeutic regimen for hemophilia B, as well as formulations that have reduced side effects and can be administered by both intravenous and subcutaneous routes.
[0149] Coagulation factors for inclusion in the BPXTEN of the present invention can include proteins of biological, therapeutic, or prophylactic benefit or function that are useful in the prevention, treatment, intervention, or amelioration of blood clotting disorders, diseases, or deficiencies. Suitable coagulation proteins include biologically active polypeptides that participate in the coagulation cascade as substrates, enzymes, or cofactors.
[0150] Table 4e provides a non-limiting list of clotting factor sequences encompassed by the BPXTEN fusion proteins of the invention. Clotting factors for inclusion in the BPXTEN of the invention can be proteins exhibiting at least about 80% sequence identity, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a protein sequence selected from Table 4e. [Table 9-1] [Table 9-2]
[0151] "Factor IX" ("FIX") includes the human factor IX protein and species and sequence variants thereof that have at least some of the biological receptor activity of mature factor IX. In some embodiments, the FIX peptide is a structural analog or peptidomimetic of any of the FIX peptides described herein, including the sequences in Table 4e. In some embodiments, the FIX peptide is a structural analog or peptidomimetic of any of the FIX peptides described herein, including the sequences in Table 4e. In one specific example of the invention, FIX is human FIX. In another embodiment, FIX is a polypeptide sequence in Table 4e. Mature factor IX is a single-chain protein of 415 amino acid residues containing approximately 17% carbohydrate by weight (Schmidt 2003, Trends Cardiovasc Med,13:39).
[0152] In some cases, the coagulation factor is factor IX, a sequence variant of factor IX, or a portion of factor IX, such as the exemplary sequences in Table 4e, and any protein or polypeptide substantially homologous thereto whose biological properties result in the activity of factor IX.
[0153] "Factor VII" (FVII) refers to the human protein, as well as species and sequence variants thereof that have at least some of the biological activity of activated factor VII. Factor VII and recombinant human FVIIa have been introduced for use in treating uncontrollable bleeding in hemophiliacs (with factor VIII or factor IX deficiency) who have developed inhibitors to the replacement clotting factor. Recombinant human factor VIIa has utility in treating uncontrollable bleeding in hemophiliacs (with factor VIII or factor IX deficiency), including those who have developed inhibitors to the replacement clotting factor. In some embodiments, the FVII peptide is an activated form (FVIIa) that is a structural analog or peptidomimetic of any of the FVII peptides described herein, including the sequences in Table 4e. Factor VII and factor VIIa are cloned as described in U.S. Patent No. 6,806,063 and U.S. Patent Application Publication No. 20080261886.
[0154] Growth hormone protein "Growth hormone" or "GH" refers to human growth hormone protein and its species and sequence variants, including, but not limited to, the 191-amino acid single-stranded human sequence of GH. The present invention contemplates the inclusion in BPXTEN of any GH-homologous sequence, including natural sequences from primates, mammals (including domestic animals), and non-natural sequence variants that retain at least some of the biological activity or function of GH, and / or sequence fragments that are useful for preventing, treating, intervening, or alleviating GH-related diseases, deficiencies, disorders, or conditions. Non-mammalian GH sequences are well described in the literature. For example, a sequence alignment of fish GH can be found in Genetics and Molecular Biology 2003, 26, pp. 295-300. In addition, natural sequences homologous to human GH can be found by standard homology search techniques, such as NCBI BLAST.
[0155] In one embodiment, the GH incorporated into a human or animal composition can be a recombinant polypeptide having a sequence corresponding to the protein found in nature. In another embodiment, the GH can be a sequence variant, fragment, homolog, or mimetic of the native sequence that retains at least some of the biological activity of native GH. Table 4f provides a non-limiting list of GH sequences from a wide variety of mammalian species encompassed by the BPXTEN fusion proteins of the invention. Any of these GH sequences or homologous derivatives constructed by recombining individual mutations between species or families can be useful in the fusion proteins of the invention. GH that can be incorporated into a BPXTEN fusion protein can include proteins that exhibit at least about 80% sequence identity, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a protein selected from Table 4f. [Table 10-1] [Table 10-2]
[0156] cytokines A BP can be a cytokine or one or more cytokines. Cytokines refer to proteins released by cells that can affect cell behavior (e.g., chemokines, interferons, lymphokines, interleukins, and tumor necrosis factors). Cytokines can be produced by a wide range of cells, including immune cells such as macrophages, B lymphocytes, T lymphocytes, and mast cells, as well as endothelial cells, fibroblasts, and various stromal cells. A given cytokine can be produced by more than one type of cell. Ignitokines can be involved in producing systemic or local immunomodulatory effects.
[0157] Certain cytokines can function as pro-inflammatory cytokines. Pro-inflammatory cytokines refer to cytokines that are involved in inducing or amplifying inflammatory responses. Pro-inflammatory cytokines can work with various cells of the immune system, such as neutrophils and leukocytes, to produce an immune response. Certain cytokines can function as anti-inflammatory cytokines. Anti-inflammatory cytokines refer to cytokines that are involved in reducing inflammatory responses. Anti-inflammatory cytokines can, in some cases, regulate pro-inflammatory cytokine responses. Some cytokines can function as both pro-inflammatory and anti-inflammatory cytokines.
[0158] Cytokines encompassed by the compositions of the present invention may have utility in the treatment of a variety of therapeutic or disease categories, including, but not limited to, cancer, rheumatoid arthritis, multiple sclerosis, myasthenia gravis, systemic lupus erythematosus, Alzheimer's disease, schizophrenia, viral infections (e.g., chronic hepatitis C, AIDS), allergic asthma, retinal neurodegenerative processes, metabolic disorders, insulin resistance, and diabetic cardiomyopathy. Cytokines may be particularly useful in the treatment of inflammatory and autoimmune conditions.
[0159] Examples of cytokines that can be regulated by the systems and compositions of the present disclosure include, but are not limited to, lymphokines, monokines, and traditional polypeptide hormones, excluding human growth hormone. Among cytokines, glycoprotein hormones such as parathyroid hormone, thyroxine, insulin, proinsulin, relaxin, prorelaxin, follicle-stimulating hormone (FSH), thyroid-stimulating hormone (TSH), and luteinizing hormone (LH), hepatic growth factor, fibroblast growth factor, prolactin, placental lactogen, tumor necrosis factor-alpha, Müllerian inhibitory factor, mouse gonadotropin-related peptide, inhibin, activin, vascular endothelial growth factor, integrins, thrombopoietin (TPO), nerve growth factor such as NGF-alpha, platelet growth factor, transforming growth factors (TGFs) such as TGF-alpha, TGF-beta, TGF-beta1, TGF-beta2, and TGF-beta3, insulin-like growth factor-I and II, erythropoietin (EPO), Flt-3L, stem cell factor (SCF), osteoinductive factor, interferons (IFNs) such as IFN-α, IFN-β, and IFN-γ, N), colony-stimulating factors (CSFs) such as macrophage-CSF (M-CSF), granulocyte-macrophage-CSF (GM-CSF), granulocyte-CSF (G-CSF), macrophage-stimulating factor (MSP), IL-1, IL-1a, IL-1b, IL-1RA, IL-18, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-9, IL-10, IL-11, IL-12, IL-12b, IL-13, IL-14, IL These include interleukins (ILs) such as IL-15, IL-16, IL-17, and IL-20; tumor necrosis factors such as CD154, LT-beta, TNF-alpha, TNF-beta, 4-1BBL, APRIL, CD70, CD153, CD178, GITRL, LIGHT, OX40L, TALL-1, TRAIL, TWEAK, and TRANCE; and other polypeptide factors including LIF, oncostatin M (OSM), and Kit ligand (KL). Cytokine receptors refer to receptor proteins that bind to cytokines. Cytokine receptors can be both membrane-bound and soluble.
[0160] The target polynucleotide may encode a cytokine. Non-limiting examples of cytokines include 4-1BBL, activin βA, activin βB, activin βC, activin βE, artemin (ARTN), BAFF / BLyS / TNFSF138, BMP10, BMP15, BMP2, BMP3, BMP4, BMP5, BMP6, BMP7, BMP8a, BMP8b, bone morphogenetic protein 1 (BMP1), CCL1 / TCA3, CCL11, CCL12 / MCP-5, CCL13 / MCP-4, CCL14, CCL15, CCL16, CCL17 / TARC, CCL18, CCL19, CCL2 / MCP-1, CCL 20, CCL21, CCL22 / MDC, CCL23, CCL24, CCL25, CCL26, CCL27, CCL28, CCL3, CCL3L3, CCL4, CCL4L1 / LAG-1, CCL5, CCL6, CCL7, CCL8, CCL9, CD153 / C D30L / TNFSF8, CD40L / CD154 / TNFSF5, CD40LG, CD70, CD70 / CD27L / TNFSF7, CLCF1, c-MPL / CD110 / TPOR, CNTF, CX3CL1, CXCL1, CXCL10, CXCL11, C XCL12, CXCL13, CXCL14, CXCL15, CXCL16, CXCL17, CXCL2 / MIP-2, CXCL3, CXCL4, CXCL5, CXCL6, CXCL7 / Ppbp, CXCL9, EDA-A1, FAM19A1, FAM19A2, FAM19A3, FAM19A4, FAM19A5, Fas ligand / FASLG / CD95L / CD178, GDF10, GDF11, GDF15, GDF2, GDF3, GDF4, GDF5, GDF6, GDF7, GDF8, GDF9, and glial cell line-derived phosphodiesterases (GDCs). Transtrophic factor (GDNF), growth differentiation factor 1 (GDF1), IFNA1, IFNA10, IFNA13, IFNA14, IFNA2, IFNA4, IFNA5 / IFNaG, IFNA7, IFNA8, IFNB1, IFNE, IFNG, IFNZ, IFNω / IFNW1, IL11, IL18, IL18BP, IL1A, IL1B, IL1F10, IL1F3 / IL1RA, IL1F5, IL1F6, IL1F7, IL1F8, IL1F9, IL1RL2, IL31, IL33, IL6, IL8 / CXCL8, inhibin-A , inhibin-B, leptin, LIF, LTA / TNFB / TNFSF1, LTB / TNFC, neurturin (NRTN), OSM, OX-40L / TNFSF4 / CD252, persephin (PSPN), RANKL / OPGL / TNFSF11 (CD254), TL1A / TNFSF15, TNFA, TNF-alpha / TNFA, TNFSF10 / TRAIL / APO-2L (CD253), TNFSF12, TNFSF13, TNFSF14 / LIGHT / CD258, XCL1, and XCL2. In some embodiments, the target gene encodes an immune checkpoint inhibitor.Non-limiting examples of such immune checkpoint inhibitors include PD-1, CTLA-4, LAG3, TIM-3, A2AR, B7-H3, B7-H4, BTLA, IDO, KIR, and VISTA. In some embodiments, the target gene encodes a T cell receptor (TCR) alpha, beta, gamma, and / or delta chain.
[0161] In some cases, the cytokine can be a chemokine, including, but not limited to, ARMCX2, BCA-1 / CXCL13, CCL11, CCL12 / MCP-5, CCL13 / MCP-4, CCL15 / MIP-5 / MIP-1 delta, CCL16 / HCC-4 / NCC4, CCL17 / TARC, CCL18 / PARC / MIP-4, CCL19 / MIP-3b, CCL2 / MCP-1, CCL20 / MIP-3 alpha / MIP3A, CCL21 / 6Ckine, CCL22 / MDC, CCL23 / MIP 3, CCL24 / eotaxin-2 / MPIF-2, CCL25 / TECK, CCL26 / eotaxin-3, CCL27 / CTACK, CCL28, CCL3 / Mip1a, CCL4 / MIP1 B, CCL4L1 / LAG-1, CCL5 / RANTES, CCL6 / C10, CCL8 / MCP-2, CCL9, CML5, CXCL1, CXCL10 / Crg-2, CXCL12 / SDF-1 beta, CXCL14 / BRAK, CXCL15 / Langmuir, CXCL16 / SR-PSOX, CXCL17, CXCL2 / MIP-2, CXCL3 / GRO gamma, CXCL4 / PF4, CXCL5, CXCL6 / GCP-2, CXCL9 / MIG, FAM19A1, FAM19A2, FAM19A3, FAM19A4 / TAFA4, FAM19A5, fractalkine / CX3CL1, I-309 / CCL1 / TCA-3, IL-8 / CXCL8, MCP-3 / CCL7, NAP-2 / PPBP / CXCL7, XCL2, and IL10.
[0162] Table 4g provides a non-limiting list of such sequences of BPs encompassed by the BPXTEN fusion proteins of the invention. 4g. The protein may be a protein exhibiting at least about 80% sequence identity, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a protein sequence selected from the group consisting of: [Table 11]
[0163] "IL-1ra" refers to human IL-1 receptor antagonist proteins and species and sequence variants, including the sequence variant anakinra (Kineret®), which have at least some of the biological activity of mature IL-1ra. Human IL-1ra is a mature glycoprotein of 152 amino acid residues. IL-1ra-containing fusion proteins of the invention may find particular use in the treatment of any of the aforementioned diseases and disorders. IL-1ra is cloned as described in U.S. Pat. Nos. 5,075,222 and 6,858,409.
[0164] In some cases, the BP may be IL-10. IL-10 may be an effective anti-inflammatory cytokine that suppresses the production of pro-inflammatory cytokines and chemokines. IL-10 may be useful in the treatment of autoimmune and inflammatory diseases such as rheumatoid arthritis, multiple sclerosis, myasthenia gravis, systemic lupus erythematosus, Alzheimer's disease, schizophrenia, allergic asthma, retinal neurodegenerative processes, and diabetes.
[0165] In some cases, IL-10 can be modified to improve stability and reduce thermal degradation. The modification can be one or more amide bond substitutions. In some cases, one or more amide bonds in the backbone of IL-10 can be substituted to achieve the above-mentioned effects. One or more amide bonds (-CONH-) in IL-10 can be substituted with Amide bonds such as -CH2NH-, -CH2S-, -CH2CH2-, -CH=CH- (cis and trans), -COCH2-, -CH(OH)CH2- or -CH2SO- In addition, the amide bond in IL-10 can also be replaced by a reduced isosteric pseudopeptide bond. See Couder et al. (1993) Int. J. Peptide Protein Res. 41:181-184, incorporated herein by reference in its entirety.
[0166] One or more acidic amino acids may be substituted, including aspartic acid, glutamic acid, homoglutamic acid, tyrosine, alkyl, aryl, arylalkyl, and heteroaryl sulfonamides of 2,4-diaminopropionic acid, ornithine or lysine, and tetrazole-substituted alkyl amino acids, as well as side chain amide residues such as asparagine, glutamine, and alkyl or aromatic-substituted derivatives of asparagine or glutamine, and serine, threonine, homoserine, 2,3-diaminopropionic acid, and alkyl or aromatic-substituted derivatives of serine or threonine.
[0167] One or more hydrophobic amino acids in IL-10, such as alanine, leucine, isoleucine, valine, norleucine, (S)-2-aminobutyric acid, (S)-cyclohexylalanine, or other simple alpha-amino acids, may be substituted with an amino acid containing an aliphatic side chain of C1 to C10 carbons, including, but not limited to, branched, cyclic, and straight chain alkyl, alkenyl, or alkynyl substitutions.
[0168] In some cases, the one or more hydrophobic amino acids in IL-10 are, for example, 2-, 3-, or 4-aminophenylalanine, 2-, 3-, or 4-chlorophenylalanine, 2-, 3-, or 4-methylphenylalanine, 2-, 3-, or 4-methoxyphenylalanine, 5-amino-, 5-chloro-, 5-methyl-, or 5-methoxytryptophan, 2'-, 3'-, or 4'-amino-, 2'-, 3'-, or 4'-chloro-, 2, 3, or 4-biphenylalanine, 2'-, 3'-, or 4'-methyl-, 2-, 3-, or The aromatic amino acids may be substituted with aromatic-substituted hydrophobic amino acid substitutions, including phenylalanine, tryptophan, tyrosine, sulfotyrosine, biphenylalanine, 1-naphthylalanine, 2-naphthylalanine, 2-benzothienylalanine, 3-benzothienylalanine, histidine, 4-biphenylalanine, and 2- or 3-pyridylalanine, including amino, alkylamino, dialkylamino, aza, halogenated (fluoro, chloro, bromo, or iodo) or alkoxy (C1-C4) substituted forms of the aromatic amino acids listed above.
[0169] One or more hydrophobic amino acids in IL-10, such as phenylalanine, tryptophan, tyrosine, sulfotyrosine, biphenylalanine, 1-naphthylalanine, 2-naphthylalanine, 2-benzothienylalanine, 3-benzothienylalanine, histidine, including amino, alkylamino, dialkylamino, aza, halogenated (fluoro, chloro, bromo, or iodo), or alkoxy, are substituted with 2-, 3-, or 4-aminophenylalanine, 2-, 3-, or 4-chlorophenylalanine. The amino acids may be substituted with aromatic amino acids including tryptophan, 2-, 3-, or 4-methylphenylalanine, 2-, 3-, or 4-methoxyphenylalanine, 5-amino-, 5-chloro-, 5-methyl-, or 5-methoxytryptophan, 2'-, 3'-, or 4'-amino-, 2'-, 3'-, or 4'-chloro-, 2-, 3, or 4-biphenylalanine, 2'-, 3'-, or 4'-methyl-, 2-, 3-, or 4-biphenylalanine, and 2- or 3-pyridylalanine.
[0170] Amino acids containing basic side chains, including arginine, lysine, histidine, ornithine, 2,3-diaminopropionic acid, homoarginine, including alkyl, alkenyl, or aryl substituted derivatives of the foregoing amino acids, can be substituted. The N-epsilon-isopropyl-lysine, 3-(4-tetrahydropyridyl)-glycine, 3-(4-tetrahydropyridyl)-alanine, N,N-gamma, gamma'-diethyl-homoarginine, alpha-methyl-arginine, alpha-methyl-2,3-diaminopropionic acid, alpha-methyl-histidine, and alpha-methyl-ornithine occupy the pro-R position of the -carbon. The modified IL-10 may contain any combination of alkyl, aromatic, heteroaromatic, ornithine, or 2,3-diaminopropionic acid, carboxylic acids, or amides formed from any of the many well-known activated derivatives, such as acid chlorides, active esters, active azolides and related derivatives, lysine, and ornithine.
[0171] In some cases, IL-10 can include one or more naturally occurring L-amino acids, synthetic L-amino acids, and / or D-enantiomers of amino acids. IL-10 polypeptides can include one or more of the following amino acids: ω-aminodecanoic acid, ω-aminotetradecanoic acid, cyclohexylalanine, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, δ-aminovaleric acid, t-butylalanine, t-butylglycine, N-methylisoleucine, phenylglycine, cyclohexylalanine, norleucine, naphthylalanine, ornithine, citrulline, 4-chlorophenylalanine, 2-fluorophenylalanine, pyridylalanine, 3-benzothienylalanine, hydroxyproline, β-alanine, o-aminobenzoic acid, m-aminobenzoic acid, p-aminobenzoic acid, m-aminomethylalanine, β-aminobenzoic acid, β ... The amino acids may include one or more of benzoic acid, 2,3-diaminopropionic acid, α-aminoisobutyric acid, N-methylglycine (sarcosine), 3-fluorophenylalanine, 4-fluorophenylalanine, penicillamine, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, β-2-thienylalanine, methionine sulfoxide, homoarginine, N-acetyllysine, 2,4-diaminobutyric acid, rho-aminophenylalanine, N-methylvaline, homocysteine, homoserine, ε-aminohexanoic acid, ω-aminohexanoic acid, ω-aminoheptanoic acid, ω-aminooctanoic acid, and 2,3-diaminobutyric acid.
[0172] IL-10 may contain a cysteine residue or cysteines that can act as a linker to another peptide via a disulfide bond or that can act to effect cyclization of the IL-10 polypeptide. Methods for introducing cysteines or cysteine analogs are known in the art; see, e.g., U.S. Pat. No. 8,067,532. IL-10 polypeptides can be cyclized. Other means of cyclization include the introduction of oxime or lanthionine linkers; see, e.g., U.S. Pat. No. 8,044,175. Any combination of amino acids (or non-amino acid moieties) that can form a cyclization bond can be used and / or introduced. The cyclization bond can be formed by combining an amino acid and a -(CH2) with a functional group that allows for the introduction of a cross-linking bond.n CO- or -(CH2) n Any combination of amino acids (with C6H4-CO-) can occur. Some examples include disulfides, -(CH2) n -disulfide mimetics such as carba bridges, thioacetals, thioether bridges (cystathionine or lanthionine) and ester and ether containing bridges.
[0173] IL-10 can be substituted with N-alkyl, aryl, or backbone crosslinked derivatives, C-terminal hydroxymethyl derivatives, ortho-modified derivatives, N-terminal modified derivatives, including lactam constructs and substituted amides such as alkylamides and hydrazides. In some cases, the IL-10 polypeptide is a retroinverso analog.
[0174] IL-10 can be the native protein, a peptide fragment IL-10, or a modified peptide that has at least some of the biological activity of native IL-10. IL-10 can be modified to improve cellular uptake. One such modification can be the attachment of a protein transduction domain. The protein transduction domain is located at the C-terminus of IL-10. The protein transduction domain may be attached to the N-terminus of IL-10. Alternatively, the protein transduction domain may be attached to the N-terminus of IL-10. The protein transduction domain may be attached to IL-10 via a covalent bond. The protein transduction domain may be selected from any of the sequences listed in Table 4h. [Table 12]
[0175] BPs of human or animal compositions are not limited to naturally occurring full-length polypeptides, but also include recombinant versions and biologically and / or pharmacologically active variants or fragments thereof. For example, one skilled in the art will understand that various amino acid substitutions can be made in a BP to create variants with respect to the biological activity or pharmacological properties of the BP without departing from the spirit of the present invention. Examples of conservative substitutions of amino acids in a polypeptide sequence are shown in Table 5. However, in embodiments of BPXTEN in which the sequence identity of the BP is less than 100% compared to the specific sequences disclosed herein, the present invention contemplates the substitution of any of the other 19 naturally occurring L-amino acids for a given amino acid residue of a given BP, which may be at any position within the sequence of the BP, including adjacent amino acid residues. If any particular substitution results in an undesired change in biological activity, then alternative amino acids can be utilized and the constructs evaluated by the methods described herein, or using any of the techniques and guidelines for conservative and non-conservative mutations described, for example, in U.S. Patent No. 5,364,934, the contents of which are incorporated by reference in their entirety, or using methods generally known to those of skill in the art. Additionally, variants can include, for example, polypeptides in which one or more amino acid residues have been added or deleted at the N- or C-terminus of the full-length native amino acid sequence of the BP that retain at least some of the biological activity of the native peptide. [Table 13]
[0176] In some embodiments, the BP incorporated into the BPXTEN polypeptide has at least about 80% sequence identity to a sequence in Tables 4a-4h, alternatively, at least about 81%, or about 82%, or about 83%, or about 84%, or about 85%, or about 86%, or about 87%, or about 88%, or about 89%, or about 90%, or about 91%, or about 92%, or about 93%, or about 94%, or about 95%, or about 96%, or about 97%, or about 98%, or may have a sequence exhibiting about 99%, or 100% sequence identity. In some embodiments, the BP incorporated into BPXTEN may be a bispecific sequence comprising a first binding domain and a second binding domain, wherein the first binding domain having specific binding affinity for a tumor-specific marker or antigen of a target cell has at least about 80% sequence identity, or alternatively, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1109%, 1110, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151, 152, 153%, 154%, 155%, 156%, 157%, 158%, 159%, 160%, 161%, 162%, 163%, 164%, 165%, 166%, 167%, 168%, The second binding domain exhibits 5%, 96%, 97%, 98%, 99%, or 100% sequence identity and has specific binding affinity for effector cells exhibits at least about 80% sequence identity, or alternatively, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, to the paired VL and VH sequences of an anti-target cell antibody selected from Table 6a. The BPs of the foregoing embodiments can be evaluated for activity using the assays or measured or determined parameters described herein, and those sequences that retain at least about 40%, or about 50%, or about 55%, or about 60%, or about 70%, or about 80%, or about 90%, or about 95% or more of the activity compared to the corresponding native BP sequence would be considered suitable for inclusion in a human or animal BPXTEN. BPs found to retain a suitable level of activity can be linked to one or more XTEN polypeptides described above or elsewhere herein. In one embodiment, BPs found to retain a suitable level of activity can be linked to one or more XTEN polypeptides having at least about 80% sequence identity (e.g., at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity) to a sequence in Tables 3a-3b, resulting in a chimeric fusion protein.
[0177] T Cell Engager An additional structural formula for BPXTEN relates to an XTENized protease-activated T cell engager ("XPAT" or "XPATs"), where BP is a bispecific antibody (e.g., a bispecific T cell engager). In some embodiments, the XPAT composition comprises a first portion comprising a first binding domain and a second binding domain, a second portion comprising a release segment, and a third portion comprising an XTEN bulking portion. In some embodiments, the XPAT composition comprises Formula Ia (shown N-terminal to C-terminal): (First part)-(Second part)-(Third part)(Ia) wherein the first portion is bispecific comprising two scFvs, the first binding domain has specific binding affinity for a tumor-specific marker or antigen of a target cell, the second binding domain has specific binding affinity for an effector cell, the second portion comprises a release segment (RS) capable of being cleaved by a mammalian protease (the protease can be tumor- or antigen-specific and thereby activating, as described more fully herein below), and the third portion is a bulking moiety. In the foregoing embodiments, the binding domains of the first portion may be in the order (VL-VH)1-(VL-VH)2 (where "1" and "2" represent the first and second binding domains, respectively), or (VL-VH)1-(VH-VL)2, or (VH-VL)1-(VL-VH)2, or (VH-VL)1-(VH-VL)2 (where the paired binding domains are linked by a polypeptide linker, described more fully herein below). In one embodiment, alternatives for the first portions VL and VH are identified in Tables 6a-6f, alternatives for RS are identified in the sequences set forth in Tables 8a-8b (described more fully herein below), and alternatives for the bulking portion are identified herein by XTEN, albumin-binding domain, albumin, IgG-binding domain, polypeptide consisting of proline, serine, and alanine, fatty acid, Fc domain, polyethylene glycol (PEG), PLGA, and hydroxyethyl starch. If desired, the bulking portion is XTEN having at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a sequence identified by a sequence set forth in Tables 3a-3b. In the foregoing embodiment, the composition is a recombinant fusion protein. In another embodiment, the portions are attached by chemical conjugation.
[0178] In another embodiment, the XPAT composition has Formula IIa (shown N-terminal to C-terminal): (Third Part)-(Second Part)-(First Part)(IIa) wherein the first portion is bispecific comprising two scFvs, the first binding domain has specific binding affinity for a tumor-specific marker or an antigen of a target cell, the second binding domain has specific binding affinity for an effector cell, the second portion comprises a release segment (RS) capable of being cleaved by a mammalian protease, and the third portion is a bulking portion. In the foregoing embodiment, the binding domains of the first portion have the following configuration: (VL-VH)1-(VL-VH)2 (where "1" and "2" represent the first and second binding domains, respectively), or (VL-VH)1-(VH-VL)2, or (VH-VL)1-(VL-VH)2, or (VH-VL)1-( VH-VL)2, where the paired binding domains are linked by a polypeptide linker as described herein below. In one embodiment, alternatives for the first portion VL and VH are identified in Tables 6a-6f, alternatives for RS are identified in the sequences set forth in Tables 8a-8b, and alternatives for the bulking portion are identified herein by XTEN, an albumin binding domain, albumin, an IgG binding domain, a polypeptide consisting of proline, serine, and alanine, a fatty acid, and an Fc domain. If desired, the bulking portion is an XTEN having at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a sequence selected from the group of sequences set forth in Tables 3a-3b. In the foregoing embodiment, the composition is a recombinant fusion protein. In another embodiment, the portions are joined by chemical conjugation.
[0179] In another embodiment, the XPAT composition has the formula IIIa (shown from N-terminus to C-terminus): (fifth moiety)-(fourth moiety)-(first moiety)-(second moiety)-(third moiety)(IIIa). wherein the first portion is bispecific comprising two scFvs, the first binding domain has specific binding affinity for a tumor-specific marker or an antigen of a target cell, the second binding domain has specific binding affinity for an effector cell, the second portion comprises a release segment (RS) capable of being cleaved by a mammalian protease, the third portion is a bulking portion, the fourth portion comprises a release segment (RS) capable of being cleaved by a mammalian protease that may be the same as or different from the second portion, and the fifth portion is a bulking portion that may be the same as or different from the third portion. In the foregoing embodiments, the binding domains of the first portion can be in the order (VL-VH)1-(VL-VH)2 (where "1" and "2" represent the first and second binding domains, respectively), or (VL-VH)1-(VH-VL)2, or (VH-VL)1-(VL-VH)2, or (VH-VL)1-(VH-VL)2 (where the paired binding domains are connected by a polypeptide linker, as described herein below). In the foregoing embodiments, alternatives to the RS are identified in the sequences set forth in Tables 8a-8b. In the foregoing embodiments, alternatives to the bulking moiety are identified herein by XTEN, an albumin binding domain, albumin, an IgG binding domain, a polypeptide consisting of proline, serine, and alanine, a fatty acid, and an Fc domain. If desired, the bulking moiety is an XTEN having at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a sequence selected from the group of sequences set forth in Tables 3a-3b. In the foregoing embodiment, the composition is a recombinant fusion protein. In another embodiment, the moieties are attached by chemical conjugation.
[0180] Based on its design and specific components, the human or animal composition advantageously provides bispecific therapeutic agents with greater selectivity, longer half-lives, which, when cleaved by proteases found in association with diseased target tissues or tissues, are less toxic and produce fewer side effects, and the human or animal composition has an improved therapeutic index compared to bispecific antibody compositions known in the art. Such compositions are useful in the treatment of certain diseases, including, but not limited to, cancers as described herein. While not wishing to be bound by any mechanism, one of skill in the art will appreciate that the compositions of the present invention achieve this reduction in nonspecific interactions through a combination of mechanisms, including steric hindrance by positioning the binding domains on bulky XTEN molecules; the flexible, amorphous character of the XTEN polypeptide, by being tethered by the composition, allows it to fluctuate and move around the binding domains, thereby resulting in a blockage between the composition and the tissue or cell, as well as the large molecular weight (both the actual molecular weight of the XTEN polypeptide and the amorphous XTEN polypeptide) compared to the size of the individual binding domains. It will be understood that this will result in a reduced ability of the intact composition to penetrate cells or tissues (due to the large hydrodynamic radius of the peptide). However, when the composition is in proximity to a target tissue or cell that produces or secretes a protease capable of cleaving the RS, or when the binding domain binds to a ligand, and is internalized into the target cell or tissue, the bispecific binding domain is liberated from the bulk of the XTEN by the action of the protease, thereby removing the steric hindrance barrier and freeing it to exert its pharmacological effect. Human or animal compositions find use in the treatment of various conditions where selective delivery to cells, tissues, or organs is desired. In one embodiment, the target tissue is cancer, which can be leukemia, lymphoma, or a tumor of an organ or system.
[0181] Binding domain The present disclosure contemplates the use of single-chain binding domains, including, but not limited to, Fv, Fab, Fab', Fab'-SH, F(ab'), linear antibodies, single-domain antibodies, single-domain camelid antibodies, single-chain antibody molecules (scFv), and diabodies capable of binding to ligands or receptors associated with effector cells and antigens of diseased tissue or cells, such as cancer, tumor, or other malignant tissue. In some embodiments, a bispecific antibody comprises a first binding domain with binding specificity for a target cell marker and a second binding domain with binding specificity for an effector cell antigen. In some embodiments, the first and second binding domains can be non-antibody scaffolds, such as anticalins, adnectins, finomers, affilins, affibodies, centrins, or DARPins. In other embodiments, the tumor cell-targeting binding domain is a variable domain of a T-cell receptor engineered to bind to MHC loaded with a peptide fragment of a protein overexpressed by tumor cells. In some embodiments, XPAT compositions are designed to provide a wide therapeutic window by taking into account the location of the target tissue protease and the presence of the same protease in healthy tissues that are not intended to be targeted, as well as the presence of the target ligand in healthy tissues but a greater presence of the ligand in unhealthy target tissues. The "therapeutic window" refers to the maximum difference between the minimum effective dose and the maximum tolerated dose for a given therapeutic composition. To help achieve a wide therapeutic window, the binding domain of the first portion of the composition is shielded by a proximal bulking moiety (e.g., an XTEN polypeptide), reducing the binding affinity of the intact composition for one or both of the ligands compared to a composition cleaved by a mammalian protease, thereby releasing the first portion from the shielding effect of the bulking moiety.
[0182] With regard to single-chain binding domains, it is well established in the art that Fvs are the smallest antibody fragments containing a complete antigen recognition and binding site, consisting of a dimer of one heavy chain (VH) and one light chain variable domain (VL) in noncovalent association. Within each VH and VL chain, there are three complementarity-determining regions (CDRs) that interact to define an antigen-binding site on the surface of the VH-VL dimer, and the six CDRs of the binding domain confer antigen-binding specificity to the antibody or single-chain binding domain. In some cases, scFvs are generated, each with three, four, or five CHRs within each binding domain. The framework sequences flanking the CDRs have a tertiary structure that is essentially conserved in native immunoglobulins across species, and the framework residues (FRs) serve to hold the CDRs in their proper orientation. The constant domains are not required for binding function but may serve to stabilize the VH-VL interaction. In some embodiments, the domains of the binding site of the polypeptide can be pairs of VH-VL, VH-VH, or VL-VL domains of either the same or different immunoglobulins, but it is generally preferred to use the respective VH and VL chains from the parent antibody to create a single-chain binding domain. The order of the VH and VL domains within the polypeptide chain is not limiting for the present invention, and the order of a given domain can usually be reversed without loss of function, but it is understood that the VH and VL domains are positioned so that the antigen-binding site can be correctly folded. Thus, the order of the VH and VL domains of the human or animal composition can be reversed. The single chain binding domains of the bispecific scFv embodiment are in the order (VL-VH) 1 -(VL-VH) 2 (where "1" and "2" represent the first and second binding domains, respectively), or (VL-VH) 1 -(VH-VL) 2 , or (VH-VL) 1 -(VL-VH) 2 , or (VH-VL) 1 -(VH-VL) 2 wherein the paired binding domains are linked by a polypeptide linker as described herein below.
[0183] Thus, the binding domain arrangement in the exemplary bispecific single-chain antibodies disclosed herein can be such that the first binding domain is C-terminal to the second binding domain. The V chain arrangement can be VH(target cell surface antigen)-VL(target cell surface antigen)-VL(effector cell antigen)-VH(effector cell antigen), VH(target cell surface antigen)-VL(target cell surface antigen)-VH(effector cell antigen)-VL(effector cell antigen), VL(target cell surface antigen)-VH(target cell surface antigen)-VL(effector cell antigen)-VH(effector cell antigen), or VL(target cell surface antigen)-VH(target cell surface antigen)-VH(effector cell antigen)-VL(effector cell antigen). The following arrangements are possible where the second binding domain is located N-terminal to the first binding domain: VH(effector cell antigen)-VL(effector cell antigen)-VL(target cell surface antigen)-VH(target cell surface antigen), VH(effector cell antigen)-VL(effector cell antigen)-VH(target cell surface antigen)-VL(target cell surface antigen), VL(effector cell antigen)-VH(effector cell antigen)-VL(target cell surface antigen)-VH(target cell surface antigen), or VL(effector cell antigen)-VH(effector cell antigen)-VH(target cell surface antigen)-VL(target cell surface antigen). As used herein, "N-terminally" or "C-terminally" and grammatical variations thereof refer to relative positions within the primary amino acid sequence rather than substitution at the absolute N- or C-terminus of the bispecific single chain antibody. Thus, as a non-limiting example, a first binding domain that is "located C-terminal to a second binding domain" indicates that the first binding domain is located on the carboxyl side of the second binding domain within the bispecific single chain antibody, and does not exclude the possibility that additional sequences, e.g., a His-tag, or another compound such as a radioisotope, are located at the C-terminus of the bispecific single chain antibody.
[0184] In one embodiment, the chimeric polypeptide assembly composition comprises a first portion comprising a first binding domain and a second binding domain, each of the binding domains being an scFv, each scFv comprising one VL and one VH. In another embodiment, the chimeric polypeptide assembly composition comprises a first portion comprising a first binding domain and a second binding domain, the binding domains being in a diabody configuration, each comprising one VL and one VH. In the foregoing embodiments, the first domain has binding specificity for a tumor-specific marker or a target cell antigen, and the second binding domain has binding specificity for an effector cell antigen. In one of the foregoing embodiments, the effector cell antigen is expressed on or within an effector cell. In one embodiment, the effector cell antigen is expressed on a T cell, such as a CD4+, CD8+, or natural killer (NK) cell. In another embodiment, the effector cell antigen is expressed on a B cell, a master cell, a dendritic cell, or a myeloid cell. In one embodiment, the effector cell antigen is CD3, a cluster of differentiation 3 antigen of cytotoxic T cells. In some of the foregoing embodiments, the first binding domain exhibits binding specificity for a tumor-specific marker associated with tumor cells. In one embodiment, the binding domain has binding affinity for a tumor-specific marker, and the tumor cells may include, but are not limited to, cells from stromal cell tumors, fibroblast tumors, myofibroblast tumors, glial cell tumors, epithelial cell tumors, adipocyte tumors, immune cell tumors, hemangioblast tumors, and smooth muscle cell tumors. In one embodiment, the tumor-specific marker or target cell antigen is alpha4 integrin, Ang2, B7-H3, B7-H6, CEACAM5, cMET, CTLA4, FOLR1, EpCAM, CCR5, CD19, HER2, HER2neu, HER3, HER4, HER1 (EGFR), or HER2-neutral tumors. FR), PD-L1, PSMA, CEA, TROP-2, MUC1 (mucin), MUC-2, MUC3, MUC4, MUC5AC, MUC5B, MUC7, MUC16 βhCG, Lewis-Y, CD20, CD33, CD38, CD30, CD56 (NCAM), CD133, ganglioside GD3; 9-O-acetyl-GD3, GM2, Globo H, fucosyl GM1, GD2, carbonic anhydrase IX, CD44v6, nectin-4, sonic hedgehog (Shh), Wue-1, plasma cell antigen 1, melanoma chondroitin sulfate proteoglycan (MCSP), CCR8, six-transmembrane epithelial antigen of the prostate (STEAP), mesothelin, A33 antigen, prostate stem cell antigen (PSCA), Ly-6, desmoglein 4, fetal acetylcholine receptor (fnAChR), CD25, cancer antigen 19-9 (CA19-9), cancer antigen 125 (CA-125), Müllerian inhibitory substance receptor type II (MISIIR), sialylated Tn antigen (sTN), The first binding domain may be fibroblast activation antigen (FAP), endosialin (CD248), epidermal growth factor receptor variant III (EGFRvIII), tumor-associated antigen L6 (TAL6), SAS, CD63, TAG72, Thomsen-Friedenreich antigen (TF-antigen), insulin-like growth factor I receptor (IGF-IR), Cora antigen, CD7, CD22, CD70, CD79a, CD79b, G250, MT-MMPs, F19 antigen, CA19-9, CA-125, alpha-fetoprotein (AFP), VEGFR1, VEGFR2, DLK1, SP17, ROR1, or EphA2. In one embodiment, the first binding domain exhibiting binding affinity for CD70 is its natural ligand, CD27, rather than an antibody fragment. In another embodiment, the first binding domain exhibiting binding affinity for B7-H6 is its natural ligand, Nkp30, rather than an antibody fragment.
[0185]
[0003] scFv embodiments of the XPAT compositions of the invention include a first binding domain and a second binding domain, wherein the VL and VH domains are derived from a monoclonal antibody having binding specificity for a tumor-specific marker or target cell antigen and an effector cell antigen, respectively. In other cases, the first and second binding domains each include six CDRs derived from a monoclonal antibody having binding specificity for a target cell marker, such as a tumor-specific marker, and an effector cell antigen, respectively. In other embodiments, the first and second binding domains of the first portion of the human or animal composition have three, four, or five CHRs within each binding domain. In other embodiments, embodiments of the invention include a first binding domain and a second binding domain, each including a CDR-H1 region, a CDR-H2 region, a CDR-H3 region, a CDR-L1 region, a CDR-L2 region, and a CDR-H3 region, wherein each of the aforementioned regions is derived from a monoclonal antibody capable of binding to a tumor-specific marker or target cell antigen and an effector cell antigen, respectively. In one embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the second binding domain comprises VH and VL regions derived from a monoclonal antibody capable of binding to human CD3. In another embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the scFv second binding domain comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity or is identical to the paired VL and VH sequences of an anti-CD3 antibody described in Table 6a. In another aspect, a second domain embodiment of the present invention comprises a CDR-H1 region, a CDR-H2 region, a CDR-H3 region, a CDR-L1 region, a CDR-L2 region, and a CDR-H3 region, wherein each of the foregoing regions is derived from a monoclonal antibody described in Table 6a. In the foregoing embodiments, the VH and / or VL domains may be configured as scFvs, diabodies, single domain antibodies, or single domain camelid antibodies.
[0186] In other embodiments, the second domain of the human or animal composition is derived from an anti-CD3 antibody listed in Table 6a. The second binding domain comprises the paired VL and VH region sequences of the anti-CD3 antibody set forth in Table 6a. In another embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the second binding domain comprises a VH and VL region, and each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity or is identical to the paired VL and VH sequence of the huUCHT1 anti-CD3 antibody set forth in Table 6a. In the foregoing embodiments, the VH and / or VL domains may be configured as part of an scFv, a diabody, a single domain antibody, or a single domain camelid antibody.
[0187] In other embodiments, the scFv of the first domain of the composition is derived from an anti-tumor cell antibody described in Table 6f. In another embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the first binding domain comprises a VH and a VL region, and each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity to the paired VL and VH sequence of an anti-tumor cell antibody described in Table 6f. In one of the foregoing embodiments, the first domain of the recited composition comprises the paired VL and VH region sequence of an anti-tumor cell antibody disclosed herein. In the foregoing embodiments, the VH and / or VL domain may be configured as an scFv, part of a diabody, a single domain antibody, or a single domain camelid antibody.
[0188] In another embodiment, a chimeric polypeptide assembly composition comprises a first portion comprising a first binding domain and a second binding domain, wherein the binding domains are in a diabody configuration, and each of the binding domains comprises one VL domain and one VH domain. In one embodiment, a diabody embodiment of the present invention comprises a first binding domain and a second binding domain, wherein the VL and VH domains are derived from a monoclonal antibody having binding specificity for a tumor-specific marker or target cell antigen and an effector cell antigen, respectively. In another embodiment, a diabody embodiment of the present invention comprises a first binding domain and a second binding domain, each comprising a CDR-H1 region, a CDR-H2 region, a CDR-H3 region, a CDR-L1 region, a CDR-L2 region, and a CDR-H3 region, wherein each of the aforementioned regions is derived from a monoclonal antibody capable of binding to a tumor-specific marker or target cell antigen and an effector cell antigen, respectively. Diabody embodiments of the present invention are envisioned to comprise a first binding domain and a second binding domain, wherein the VL and VH domains are derived from a monoclonal antibody having binding specificity for a tumor-specific marker or target cell antigen, and an effector cell antigen, respectively. In another aspect, diabody embodiments of the present invention comprise a first binding domain and a second binding domain, each comprising a CDR-H1 region, a CDR-H2 region, a CDR-H3 region, a CDR-L1 region, a CDR-L2 region, and a CDR-H3 region, each of the foregoing regions being derived from a monoclonal antibody capable of binding to a tumor-specific marker or target cell antigen, and an effector cell antigen, respectively. In one embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the diabody second binding domain comprises paired VH and VL regions derived from a monoclonal antibody capable of binding to human CD3.In another embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the diabody second binding domain comprises a VH and a VL region, and each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity or is identical to the paired VL and VH sequences of an anti-CD3 antibody listed in Table 6a. In another embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the diabody second binding domain comprises a VH and a VL region, and each VH and VL region is listed in Table 6a. The diabody first binding domain of the composition exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity to or is identical to the VL and VH sequences of the huUCHT1 antibody described herein. In other embodiments, the diabody second domain of the composition is derived from an anti-CD3 antibody described herein. In another embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the diabody first binding domain comprises a VH and a VL region, and each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity to the VL and VH sequences of an anti-tumor cell antibody described in Table 6f. In other embodiments, the diabody first domain of the composition is derived from an anti-tumor cell antibody described herein.
[0189] For human or animal compositions, therapeutic monoclonal antibodies from which VL and VH and CDR domains can be derived are known in the art.The sequences of the above-mentioned antibodies can be obtained from publicly available databases, patents, or literature references.In addition, non-limiting examples of monoclonal antibodies and VH and VL sequences from anti-CD3 antibodies are listed in Table 6a, and non-limiting examples of monoclonal antibodies and VH and VL sequences from cancer, tumor, or target cell markers are listed in Table 6f.
[0190] Anti-CD3 binding domain In some embodiments, the present invention provides chimeric polypeptide assembly compositions comprising a first portion binding domain having binding affinity for T cells. In one embodiment, the second portion binding domain comprises a VL and a VH derived from a monoclonal antibody directed against an antigen of CD3. In another embodiment, the binding domain comprises a VL and a VH derived from a monoclonal antibody directed against CD3 epsilon and CD3 delta. Monoclonal antibodies directed against CD3neu are known in the art. Illustrative, non-limiting examples of VL and VH sequences of monoclonal antibodies directed against CD3 are listed in Table 6a. In one embodiment, the present invention provides chimeric polypeptide assemblies comprising a binding domain having binding affinity for CD3 comprising the anti-CD3 VL and VH sequences listed in Table 6a. In another embodiment, the present invention provides chimeric polypeptide assemblies comprising a first portion binding domain having binding affinity for CD3 epsilon comprising the anti-CD3 epsilon VL and VH sequences listed in Table 6a. In another embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the scFv second binding domain of the first portion comprises a VH and a VL region, and each VH and VL region exhibits or is identical to at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity with or is identical to the paired VL and VH sequence of the huUCHT1 anti-CD3 antibody listed in Table 6a. In another embodiment, the present invention provides a chimeric polypeptide assembly composition comprising a binding domain with binding affinity for CD3, comprising a CDR-L1 region, a CDR-L2 region, a CDR-L3 region, a CDR-H1 region, a CDR-H2 region, and a CDR-H3 region, each derived from a respective anti-CD3 VL and VH sequence listed in Table 6a.In another embodiment, the present invention provides a chimeric polypeptide assembly composition comprising a binding domain having binding affinity for CD3, comprising a CDR-L1 region, a CDR-L2 region, a CDR-L3 region, a CDR-H1 region, a CDR-H2 region, and a CDR-H3 region, wherein the CDR sequences are RASQDIRNYLN (SEQ ID NO: 8034), YTSRLES (SEQ ID NO: 8035), QQGNTLPWT (SEQ ID NO: 8036), GYSFTGYTMN (SEQ ID NO: 8037), LINPYKGVST (SEQ ID NO: 8038), and SGYYGDSDWYFDV (SEQ ID NO: 8039).
[0191] The CD3 complex associates with the T cell antigen receptor (TCR) and regulates cell surface expression of the TCR. Peptides: A group of cell surface molecules that function in the signal transduction cascade that occurs when an MHC ligand binds to the TCR. Typically, when an antigen binds to the T cell receptor, CD3 sends a signal through the cell membrane to the cytoplasm inside the T cell. This triggers activation of the T cell, which rapidly divides to generate new T cells primed to attack the specific antigen exposed to the TCR. The CD3 complex consists of the CD3 epsilon molecule along with four other membrane-bound polypeptides (CD3-gamma, delta, zeta, and beta). In humans, CD3-epsilon is encoded by the CD3E gene on chromosome 11. The intracellular domain of each CD3 chain contains an immunoreceptor tyrosine-based activation motif (ITAM), which serves as a nucleation point for the intracellular signaling machinery upon T cell receptor engagement.
[0192] Numerous therapeutic strategies modulate T cell immunity by targeting TCR signaling, particularly anti-human CD3 monoclonal antibodies (mAbs), which are widely used clinically in immunosuppressive regimens. The CD3-specific murine mAb OKT3 was the first mAb licensed for human use (Sgro, C. Side-effects of a monoclonal antibody, muromonab CD3 / orthoclone OKT3: bibliographic review. Toxicology 105:23-29, 1995), and is widely used clinically as an immunosuppressant in transplantation, type 1 diabetes, and psoriasis (Chatenoud, Clin. Transplant 7:422-430, (1993); Chatenoud, Nat. Rev. Immunol. 3:123-132 (2003); Kumar, Transplant. Proc. 30:1351-1352 (1998)). Importantly, anti-CD3 mAbs can induce partial T cell signaling and clonal paralysis (Smith, J.A., Nonmitogenic Anti-CD3 Monoclonal Antibodies Deliver a Partial T Cell Receptor Signal and Induce Clonal Anergy J. Exp. Med. 185:1413-1422 (1997)). OKT3 has been described in the literature as a T cell mitogen and a potent T cell killer (Wong, JT. The mechanism of anti-CD3 monoclonal antibodies. Mediation of cytolysis by inter-T cell bridging. Transplantation 50:683-689 (1990)). Notably, Wong's work demonstrated that target killing can be achieved by crosslinking CD3 T cells and target cells, and that neither FcR-mediated ADCC nor complementary chain immobilization is required for bivalent anti-CD3 MABs to lyse target cells.
[0193] OKT3 exhibits both mitogenic and T-cell killing activity in a time-dependent manner: following early activation of T cells resulting in cytokine release, after further administration, OKT3 blocks all known T-cell functions. It is this subsequent blockade of T-cell function that has caused OKT3 to find such widespread use as an immunosuppressant in therapeutic regimens to reduce or even abolish tissue rejection of allografts. Other antibodies specific for the CD3 molecule are disclosed in Tunnacliffe, Int. Immunol. 1 (1989), 546-50, WO 2007 / 033230 describes an anti-human monoclonal CD3 epsilon antibody, U.S. Pat. No. 5,821,337 describes the VL and VH sequences of the murine anti-CD3 monoclonal Ab UCHT1 (muxCD3, Shalaby et al., J. Exp. Med. 175, 217-225 (1992)), as well as a humanized variant of this antibody (hu UCHT1), and U.S. Patent Application No. 20120034228 discloses binding domains capable of binding to epitopes of the human and non-chimpanzee primate CD3 epsilon chain. [Table 14-1] [Table 14-2]
[0194] CD3 cell antigen-binding fragment In another aspect, the present disclosure relates to an antigen-binding fragment (AF2) having specific binding affinity for an effector cell antigen, which can be incorporated into any of the human or animal composition embodiments described herein. In some cases, the effector cell antigen is expressed on the surface of an effector cell selected from plasma cells, T cells, B cells, cytokine-induced killer cells (CIK cells), mast cells, dendritic cells, regulatory T cells (RegT cells), helper T cells, myeloid cells, and NK cells.
[0195] Various AF2s that bind to effector cell antigens are particularly useful for pairing with antigen-binding fragments that have binding affinity for the EGFR antigen associated with diseased cells or tissues in a format that results in cell killing of the diseased cells or tissues. Binding specificity can be determined by complementarity-determining regions, or CDRs, such as light chain CDRs or heavy chain CDRs. In this case, the binding specificity is determined by the light chain CDR and the heavy chain CDR. A given combination of heavy chain CDR and light chain CDR provides a given binding pocket that confers higher affinity and / or specificity for an effector cell antigen compared to other reference antigens. The resulting bispecific composition, in which a first antigen-binding fragment (AF1) against EGFR is linked to a second antigen-binding fragment (AF2) with binding specificity for an effector cell antigen via a short, flexible peptide linker, is a bispecific in which each antigen-binding fragment has specific binding affinity for its respective ligand. Those skilled in the art will understand that in such a composition, AF1 directed against EGFR of the diseased tissue is used in combination with AF2 directed toward an effector cell marker to bring effector cells into close proximity with the cells of the diseased tissue, resulting in cytolysis of the cells of the diseased tissue. Additionally, AF1 and AF2 are incorporated into specifically designed polypeptides comprising a cleavable release segment and XTEN to confer prodrug properties to the composition that are activated by release of the fused AF1 and AF2 upon cleavage of the release segment when in the vicinity of diseased tissue that has a protease capable of cleaving the release segment at one or more positions in the release segment sequence.
[0196] In one embodiment, the AF2 of the human or animal composition has binding affinity for an effector cell antigen expressed on the surface of T cells. In another embodiment, the AF2 of the human or animal composition has binding affinity for CD3. In another embodiment, the AF2 of the human or animal composition has binding affinity for members of the CD3 complex, including all known CD3 subunits of the CD3 complex, such as CD3 epsilon, CD3 delta, CD3 gamma, CD3 zeta, CD3 alpha, and CD3 beta, either individually or in combination. In another embodiment, the AF2 has binding affinity for CD3 epsilon, CD3 delta, CD3 gamma, CD3 zeta, CD3 alpha, or CD3 beta.
[0197] The antigen-binding fragments contemplated by the present disclosure can be derived from naturally occurring antibodies or fragments thereof, non-naturally occurring antibodies or fragments thereof, humanized antibodies or fragments thereof, synthetic antibodies or fragments thereof, hybrid antibodies or fragments thereof, or engineered antibodies or fragments thereof. Methods for generating antibodies against a given target marker are well known in the art. For example, monoclonal antibodies can be produced using the hybridoma method first described by Kohler et al., Nature, 256:495 (1975), or can be produced by recombinant DNA methods (U.S. Patent No. 4,816,567). The structure of antibodies and their fragments, the variable regions of heavy and light chains (VH and VL), single-chain variable regions (scFv), complementarity-determining regions (CDRs), and domain antibodies (dAbs) are well understood. Methods for generating polypeptides having desired antigen-binding fragments that have binding affinity for a given antigen are known in the art.
[0198] Those skilled in the art will understand that the use of the term "antigen-binding fragment" with respect to the composition embodiments disclosed herein is intended to include portions or fragments of antibodies that retain the ability to bind to an antigen that is the ligand of the corresponding intact antibody. In such embodiments, the antigen-binding fragment may be, but is not limited to, a CDR and intervening framework region, a variable or hypervariable region of the antibody's light and / or heavy chain (VL, VH), a variable fragment (Fv), a Fab' fragment, a F(ab')2 fragment, a Fab fragment, a single-chain antibody (scAb), a VHH camelid antibody, a single-chain variable fragment (scFv), a linear antibody, a single-domain antibody, a complementarity-determining region (CDR), a domain antibody (dAb), a BHH-type or BNAR-type single-domain heavy chain immunoglobulin, a single-domain light chain immunoglobulin, or other polypeptides known in the art that contain fragments of antibodies capable of binding to an antigen. An antigen-binding fragment having a CDR-H and a CDR-L comprises, from the N-terminus to the C-terminus, (CDR-H)-( The VL and VH of two antigen-binding fragments can also be configured in a single-chain diabody configuration, i.e., the VL and VH of AF1 and AF2 are configured with a linker of appropriate length to allow for configuration as a diabody.
[0199] The various CD3-binding AF2s disclosed herein have been specifically modified to enhance their stability in the polypeptide embodiments described herein. Protein aggregation of antibodies continues to pose a significant challenge to their developability and remains a key area of focus in antibody production. Antibody aggregation can be triggered by partial unfolding of their domains, leading to monomer-monomer association followed by nucleation and aggregate growth. While the aggregation propensity of antibodies and antibody-based proteins can be influenced by external experimental conditions, it is strongly dependent on intrinsic antibody properties determined by their sequence and structure. While it is well known that proteins are marginally stable in their folded state, it is less understood that most proteins are inherently prone to aggregation in their unfolded or partially unfolded state, and that the resulting aggregates can be highly stable and long-lived. Reduced aggregation propensity has also been shown to be accompanied by increased expression titers, indicating that reduced protein aggregation can be beneficial throughout the development process and lead to a more efficient path to clinical trials. For therapeutic proteins, aggregates are a significant risk factor for adverse immune reactions in patients and can occur through a variety of mechanisms. Controlling aggregation can improve protein stability, manufacturability, wear rate, safety, formulation, potency, immunogenicity, and solubility. Intrinsic protein properties, such as size, hydrophobicity, electrostatics, and charge distribution, play an important role in protein solubility. It has been shown that low solubility of therapeutic proteins due to surface hydrophobicity can make formulation development more challenging and lead to poor in vivo biodistribution, undesirable pharmacokinetic behavior, and immunogenicity. Reducing the overall surface hydrophobicity of candidate monoclonal antibodies can also provide benefits and cost savings related to purification and administration regimens. Individual amino acids can be identified through structural analysis as contributing to antibody aggregation potential and can be located in both CDRs and framework regions. Residues, in particular, can be predicted to be at high risk of causing hydrophobicity problems for a given antibody.In one embodiment, the present disclosure provides an AF2 having the ability to specifically bind to CD3, wherein AF2 has at least one amino acid substitution in a framework region with a hydrophobic amino acid selected from isoleucine, leucine, or methionine relative to a parent antibody or antibody fragment. In another embodiment, CD3 AF2 has at least two amino acid substitutions in one or more framework regions with a hydrophobic amino acid selected from isoleucine, leucine, or methionine.
[0200] Changes in the net charge of a polypeptide are considered in the design of the AF2 sequence of the embodiments described herein, particularly with respect to antibodies or antibody fragments comprising certain embodiments of the invention described herein, with individual amino acid substitutions being made relative to the parent antibody used as a starting point. Related to these design considerations is the isoelectric point (pI) of the polypeptide, which is the pH at which the antibody or antibody fragment bears no net charge. Antibodies or antibody fragments typically have a net positive charge, which tends to correlate with increased blood clearance and increased tissue retention, and generally have a short half-life, while a net negative charge results in decreased tissue uptake and a longer half-life. This charge relative to framework residues can be manipulated through mutation. The isoelectric point of a polypeptide can be determined mathematically (e.g., computationally) or experimentally in in vitro assays. In some embodiments, the isoelectric points of AF1 and AF2 are designed to be within a certain range of each other, thereby promoting stability.
[0201] In one embodiment, the present disclosure provides AF2 for use in any of the polypeptide embodiments described herein comprising a CDR-L and a CDR-H, wherein AF2 (a) specifically binds to the cluster of differentiation 3 T-cell receptor (CD3), and (b) comprises CDR-H1, CDR-H2, and CDR-H3 having the amino acid sequences of SEQ ID NOs: 742, 743, and 744, respectively. In another embodiment, the disclosure provides AF2 for use in any of the polypeptide embodiments described herein, comprising a CDR-L and a CDR-H, wherein AF2 (a) specifically binds to the cluster of differentiation 3 T-cell receptor (CD3); (b) comprises CDR-H1, CDR-H2, and CDR-H3 having the amino acid sequences of SEQ ID NOs: 742, 743, and 744, respectively; and (c) comprises a CDR-L, wherein the CDR-L comprises CDR-L1 having the amino acid sequence of SEQ ID NO: 735 or 736, CDR-L2 having the amino acid sequence of SEQ ID NO: 738 or 739, and CDR-L3 having the amino acid sequence of SEQ ID NO: 740.In another embodiment, the aforementioned AF2 embodiment of this paragraph further comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises a FR-L1 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 746 and a FR-L1 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 747. FR-L2 exhibits 94%, 95%, 96%, 97%, 98%, 99% sequence identity or is identical thereto; FR-L3 exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or is identical thereto with the amino acid sequence of any one of SEQ ID NOs: 748 to 751; and FR-L4 exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or is identical thereto with the amino acid sequence of SEQ ID NO: 754. FR-L4 exhibiting 7%, 98%, 99% sequence identity or being identical thereto; FR-H1 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto with the amino acid sequence of SEQ ID NO: 755 or SEQ ID NO: 756; and FR-H1 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with the amino acid sequence of SEQ ID NO: 759. FR-H2 exhibiting or being identical to the amino acid sequence of SEQ ID NO: 760; FR-H3 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or being identical to the amino acid sequence of SEQ ID NO: 760; and FR-H4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or being identical to the amino acid sequence of SEQ ID NO: 764.In another embodiment, AF2 for use in any of the polypeptide embodiments described herein comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H). AF2 comprises FR-L1 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or is identical to the amino acid sequence of SEQ ID NO: 746; FR-L2 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or is identical to the amino acid sequence of SEQ ID NO: 747; and FR-L3 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or is identical to the amino acid sequence of SEQ ID NO: 748. FR-L3 exhibiting 7%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; FR-L4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; and FR-L5 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; and FR-L6 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94% sequence identity or being identical thereto; FR-H1 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO: 759; FR-H2 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO: 760; and FR-H4, which exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 764.In another embodiment, AF2 for use in any of the polypeptide embodiments described herein comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises FR-L1 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 746 and a heavy chain framework region (FR-H) that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 747. FR-L2 exhibiting 9%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; FR-L3 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; and FR-L4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; and FR-L5 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; FR-L4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or identical to the amino acid sequence of SEQ ID NO: 755; FR-H1 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or identical to the amino acid sequence of SEQ ID NO: 759. FR-H2 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or identical to the amino acid sequence of SEQ ID NO: 760; and FR-H4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or identical to the amino acid sequence of SEQ ID NO: 764.In another embodiment, AF2 of the human or animal polypeptide embodiments described herein comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises FR-L1 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 746; FR-L2 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 747; and FR-L3 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 750. FR-L3 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO: 754; FR-L4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO: 755; FR-H1 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with, or being identical to, the amino acid sequence of SEQ ID NO: 759. FR-H2 exhibiting or being identical to the amino acid sequence of SEQ ID NO: 760; FR-H3 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or being identical to the amino acid sequence of SEQ ID NO: 760; and FR-H4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or being identical to the amino acid sequence of SEQ ID NO: 764. In another embodiment, AF2 of the human or animal polypeptide embodiments described herein comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises a FR-L1 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 746 and a FR-L1 that exhibits at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 747. FR-L2 exhibiting 1%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; FR-L3 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; and FR-L4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or being identical thereto; and FR-L5 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96% sequence identity or being identical thereto; FR-L4 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or identical to the amino acid sequence of SEQ ID NO: 756; FR-H1 exhibiting at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or identical to the amino acid sequence of SEQ ID NO: 759. FR-H2 that is at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 760; and FR-H4 that is at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to or is identical to the amino acid sequence of SEQ ID NO: 764.
[0202] In another embodiment, the disclosure provides AF2 for use in any of the polypeptide embodiments described herein, wherein AF2 comprises a variable heavy (VH) amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to, or identical to, the amino acid sequence of SEQ ID NO: 766 or SEQ ID NO: 769. In another embodiment, the disclosure provides AF2 for use in any of the polypeptide embodiments described herein, wherein AF2 comprises a variable light (VL) amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to, or identical to, the amino acid sequence of any one of SEQ ID NOs: 765, 767, 768, 770, or 771. In another embodiment, the disclosure provides AF2 for use in any of the polypeptide embodiments described herein, wherein AF2 comprises a variable heavy (VH) amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to, or identical to, the amino acid sequence of SEQ ID NO: 766 or SEQ ID NO: 769, and a variable light (VL) amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to, or identical to, the amino acid sequence of any one of SEQ ID NOs: 765, 767, 768, 770, or 771.
[0203] In another embodiment, the present disclosure provides any of the polypeptide embodiments described herein. The present invention provides AF2 for use in the method of the present invention, wherein AF2 comprises an amino acid sequence that has at least 95%, 96%, 97%, 98%, 99% sequence identity to, or is identical to, the amino acid sequence of any one of SEQ ID NOs: 776-780.
[0204] In another aspect, the present disclosure provides AF2 antigen-binding fragments that bind to the CD3 protein complex and have enhanced stability compared to CD3-binding antibodies or antigen-binding fragments known in the art. Additionally, the CD3 antigen-binding fragments of the present disclosure are designed to confer greater stability to chimeric bispecific antigen-binding fragment compositions into which they are incorporated, resulting in improved expression and recovery, increased shelf life, and enhanced stability of the fusion protein when administered to humans or animals. In one approach, the CD3 AF2 of the present disclosure is designed to have greater thermal stability compared to certain CD3-binding antibodies and antigen-binding fragments known in the art. As a result, CD3 AF2s utilized as components of chimeric bispecific antigen-binding fragment compositions into which they are incorporated exhibit favorable pharmaceutical properties, including high thermal stability and low aggregation tendency, resulting in improved expression and recovery during manufacturing and storage, and promoting a long serum half-life. Biophysical properties such as thermal stability are often limited by antibody variable domains, whose intrinsic properties vary widely. High thermal stability is often associated with high expression levels and other desirable properties, such as low susceptibility to aggregation (Buchanan A, et al. Engineering a therapeutic IgG molecule to address cysteinylation, aggregation and enhance thermal stability and expression. MAbs 2013;5:255). Thermal stability is measured by the "melting temperature" (T), defined as the temperature at which half of the molecule is denatured. m The melting temperature of each heterodimer indicates its thermal stability. mIn vitro assays for determining the melting point of a heterodimer are known in the art. The melting point of a heterodimer can be measured using techniques such as differential scanning calorimetry (Chen et al. (2003) Pharm Res 20:1952-60, Ghirlando et al. (1999) Immunol Lett 68:47-52). Alternatively, the thermal stability of a heterodimer can be measured using circular dichroism (Murray et al. (2002) J. Chromatogr Sci 40:343-9) or as described in the Examples below.
[0205] Thermal denaturation curves of the CD3-binding fragments of the disclosure and anti-CD3 bispecific antibodies comprising the anti-CD3 binding fragments and a reference binding fragment show that the constructs of the disclosure are more resistant to thermal denaturation than an antigen-binding fragment consisting of the sequence set forth in SEQ ID NO: 781 or a control bispecific antibody comprising SEQ ID NO: 781 and a reference antigen-binding fragment that binds to an EGFR embodiment described herein. In one embodiment, the polypeptide of any of the human or animal composition embodiments described herein comprises an anti-CD3 AF2 of the embodiments described herein, wherein the T of AF2 is m is the T of an antigen-binding fragment consisting of the sequence of SEQ ID NO: 781, as determined by an increase in melting temperature in an in vitro assay. m at least 2°C higher, or at least 3°C higher, or at least 4°C higher, or at least 5°C higher, or at least 6°C higher, or at least 7°C higher, or at least 8°C higher, or at least 9°C higher, or at least 10°C higher.
[0206] In another embodiment, the polypeptide of any of the human or animal composition embodiments described herein has a dissociation constant (K) of about 10 nM to about 400 nM, or about 50 nM to about 350 nM, or about 100 nM to 300 nM, as determined in an in vitro antigen binding assay involving human or cyno CD3 antigen. dIn another embodiment, the polypeptide of any of the human or animal composition embodiments described herein comprises AF2 that specifically binds to human or cyno CD3 with a constant AF2 as determined in an in vitro antigen binding assay. When the ATP-binding protein is soluble in water, it has a dissociation constant (K) weaker than about 10 nM, or about 50 nM, or about 100 nM, or about 150 nM, or about 200 nM, or about 250 nM, or about 300 nM, or about 350 nM, or about 400 nM. d ) containing AF2, which specifically binds to human or cyno CD3. For clarity, a K of 400 d An antigen-binding fragment having a K d In another embodiment, the polypeptides of any of the human or animal composition embodiments described herein bind to their ligands more weakly than those having a respective dissociation constant (K d ) with a binding affinity that is at least 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, or at least 10-fold weaker than an antibody-binding fragment consisting of the amino acid sequence of SEQ ID NO: 781. In another embodiment, the present disclosure provides AF2 that specifically binds to CD3. dThe present invention provides a bispecific polypeptide comprising an AF2 that exhibits a binding affinity for CD3 that is at least 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 50-fold, 100-fold, or at least 1000-fold weaker than the binding affinity of an AF1 EGFR embodiment incorporated into the human or animal polypeptide, as determined by ELISA. The binding affinity of the human or animal composition for the target ligand can be assayed using a binding or competitive binding assay, such as a Biacore assay using a chip-bound receptor or binding protein, or an ELISA assay as described in U.S. Pat. No. 5,534,617, an assay described in the Examples herein, a radioreceptor assay, or other assays known in the art. The binding affinity constant can then be determined using standard methods, such as Scatchard analysis as described by van Zoelen, et al., Trends Pharmacol Sciences (1998) 19) 12):487, or other methods known in the art.
[0207] In a related aspect, the present disclosure provides an AF2 incorporated into a chimeric bispecific polypeptide composition that binds to CD3 and is designed to have an isoelectric point (pI) that confers enhanced stability to the disclosed composition relative to corresponding compositions comprising CD3-binding antibodies or antigen-binding fragments known in the art. In one embodiment, the polypeptide of any of the human or animal composition embodiments described herein comprises an AF2 that binds to CD3, wherein the AF2 exhibits a pI that is between 6.0 and 6.6, inclusive. In another embodiment, the polypeptide of any of the human or animal composition embodiments described herein comprises an AF2 that binds to CD3, wherein the AF2 exhibits a pI that is at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0 pH units lower than the pI of a reference antigen-binding fragment consisting of the sequence set forth in SEQ ID NO:781. In another embodiment, a polypeptide of any of the human or animal composition embodiments described herein comprises a CD3-binding AF2 fused to an AF1 that binds an EGFR antigen, wherein AF2 exhibits a pI that is within at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, or 1.5 pH units of the pI of AF1 that binds the EGFR antigen or an epitope thereof. In another embodiment, a polypeptide of any of the human or animal composition embodiments described herein comprises a CD3-binding AF2 fused to an AF1 that binds an EGFR antigen, wherein AF2 exhibits a pI that is within at least about 0.1 to about 1.5, or at least about 0.3 to about 1.2, or at least about 0.5 to about 1.0, or at least about 0.7 to about 0.9 pH units of the pI of AF1. It is specifically intended that by designing the pI of these two antigen-binding fragments to be within such ranges, the resulting fused antigen-binding fragments will confer a greater degree of stability to the chimeric bispecific antigen-binding fragment composition into which they are incorporated, resulting in improved expression and enhanced recovery of the fusion protein in a soluble, non-aggregated form, increased shelf life of the formulated chimeric bispecific polypeptide composition, and enhanced stability when the composition is administered to a human or animal.Separately, having AF2 and AF1 within a relatively narrow pI range allows for buffers in which both AF2 and AF1 are stable. or other solutions may be selected, thereby promoting the overall stability of the composition.
[0208] In certain embodiments, the VL and VH of the antigen-binding fragment are fused by a relatively long linker comprising 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 hydrophilic amino acids that have flexible properties when linked together. In one embodiment, the VL and VH of any of the scFv embodiments described herein have the sequence GSGEGSEGEGGGEGSEGEGSGEGGEGEGSG (SEQ ID NO: 8058). TGSGEGSEGEGGGEGSEGEGSGEGGEGEGSGT (SEQ ID NO: 8059), GATPPETGAETESPGETTGGSAESEPPGEG (SEQ ID NO: 8060), or AF1 and AF2 are linked together by a relatively long linker of hydrophilic amino acids, which is GSAAPTAGTTPSASPAPPTGGSSAAGSPST (SEQ ID NO: 8061). In another embodiment, AF1 and AF2 are linked together by a short linker of hydrophilic amino acids having 3, 4, 5, 6, or 7 amino acids. In one embodiment, the short linker sequence is the sequence SGGGGS (SEQ ID NO: 8062), GGGGS (SEQ ID NO: 8063), GGSGGS (SEQ ID NO: 8064), GGS, or GSP. In another embodiment, the disclosure provides a composition comprising a single-chain diabody in which, after folding, the first domain (VL or VH) pairs with the last domain (VH or VL) to form one scFv, and these two domains pair in the middle to form the other scFv, wherein the first and second domains and the third and last domains are fused together by one of the aforementioned short linkers, and the second and third variable domains are fused by one of the aforementioned longer linkers. As will be appreciated by those skilled in the art, the selection of short and relatively long linkers is to prevent mispairing of adjacent variable domains, thereby promoting the formation of single-chain diabody configurations comprising the VL and VH of the first and second antigen-binding fragments. [Table 15] [Table 16] [Table 17] [Table 18-1] [Table 18-2]
[0209] Anti-EpCAM binding domain In some embodiments, the present invention provides a chimeric polypeptide assembly composition comprising a binding domain having binding affinity for the tumor-specific marker EpCAM. In one embodiment, the binding domain comprises a VL and a VH derived from a monoclonal antibody against EpCAM. Monoclonal antibodies against EpCAM are known in the art. Illustrative, non-limiting examples of EpCAM monoclonal antibodies and their VL and VH sequences are listed in Table 6f. In one embodiment, the present invention provides a chimeric polypeptide assembly comprising a binding domain having binding affinity for the tumor-specific marker EpCAM comprising the anti-EpCAM VL and VH sequences listed in Table 6f. In another embodiment, the present invention provides a chimeric polypeptide assembly composition, wherein the first binding domain of the first portion comprises a VH and a VL region, and each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity with the paired VL and VH sequence of the 4D5MUCB anti-EpCAM antibody listed in Table 6f. In another embodiment, the present invention provides a chimeric polypeptide assembly composition, comprising a binding domain with binding affinity for a tumor-specific marker, comprising a CDR-L1 region, a CDR-L2 region, a CDR-L3 region, a CDR-H1 region, a CDR-H2 region, and a CDR-H3 region, each derived from the respective VL and VH sequence listed in Table 6f. [Table 19-1] [Table 19-2] [Table 19-3] [Table 19-4] [Table 19-5] Table 19-6 Table 19-7 Table 19-8 Table 19-9 Table 19-10 Table 19-11 Table 19-12 Table 19-13 Table 19-14 Table 19-15 Table 19-16 Table 19-17
[0210] Epithelial cell adhesion molecule (EpCAM, also known as the 17-1A antigen) is a 40 kDa membrane-integrated glycoprotein of 314 amino acids expressed in certain epithelia and many human carcinomas (see Balzar, The biology of the 17-1A antigen (Ep-CAM), J. Mol. Med. 1999, 77:699-712). Due to their epithelial cell origin, tumor cells from most carcinomas are expressed on primary, metastatic, and disseminated non-small cell lung carcinoma cells (Passlick, B., et al. The 17-1A antigen is expressed on primary, metastatic, and disseminated non-small cell lung carcinoma cells. Int. J. Cancer 87(4):548-552, 2000), gastric and gastroesophageal junction adenocarcinomas (Martin, I. G., Expression of the 17-1A antigen in gastric and gastroesophageal junction adenocarcinomas: a potential immunotherapeutic target? J Clin Pathol 1999;52:701-704), and breast and colorectal carcinomas (Packeisen J, et al. Detection of surface antigen 17-1A in breast and colorectal cancer. Hybridoma. 1999). 18(1):37-40), express EpCAM on their surface (more than normal healthy cells), and overexpression of EpCAM on tumor cells is a predictor of survival (Gastl, Lancet. 2000, 356, 1981-1982). Due to their epithelial cell origin, tumor cells derived from most carcinomas express EpCAM on their surface.
[0211] In one embodiment, provided herein is a bispecific chimeric polypeptide assembly composition having a first portion having a binding domain specific for EpCAM and a binding domain specific for CD3. The technical problem to be solved was to provide means and methods for the production of improved compositions exhibiting well-tolerated and more convenient pharmaceutical (less frequent administration) properties for the effective treatment and / or palliation of neoplastic diseases. The solution to the aforementioned technical problem is achieved by the embodiments disclosed herein and is characterized in the claims.
[0212] Thus, in some embodiments, the present invention relates to a chimeric polypeptide assembly composition, said composition comprising a first portion comprising a bispecific single chain antibody composition comprising at least two binding domains, one of said domains binding to an effector cell antigen, such as CD3 antigen, and a second domain binding to an EpCAM antigen, said binding domains comprising a VL and VH specific for EpCAM and a VL and VH specific for human CD3 antigen. Preferably, in embodiments, said binding domain specific for EpCAM has a VL and VH specific for 10 or more of said EpCAM-specific VL and VH as determined in an in vitro binding assay. -7 ~10 -10 K less than M d In one of the foregoing embodiments, the binding domain is in scFv format. In another of the foregoing embodiments, The binding domain is in a single chain diabody format.
[0213] In some embodiments, the present invention provides chimeric polypeptide assembly compositions comprising a first binding domain having binding affinity for a tumor-specific marker and a second binding domain that binds to an effector cell antigen, such as a CD3 antigen. Tumor-specific markers comprising these embodiments of the present invention include CCR5, CD19, HER-2, HER-3, HER-4, EGFR, PSMA, CEA, MUC1, MUC2, MUC3, MUC4, MUC5AC, MUC5B, MUC7, βhCG, Lewis-Y, CD-20, CD33, CD30, ganglioside GD3, 9-O-acetyl-GD3, Globo H, fucosyl GM1, GD-2, carbonic anhydrase IX, CD44v6, sonic hedgehog, Wue-1, plasma cell antigen 1, melanoma chondroitin sulfate proteoglycan, CCR8, six-transmembrane epithelial antigen of the prostate (STEAP), mesothelin, A33 antigen, prostate stem cell antigen (PSCA), LY-6, SAS, desmoglein 4, fetal acetylcholine receptor, CD-25, cancer antigen 19-9 (CA19-9), and the like. 19-9), cancer antigen 125 (CA-125), Müllerian inhibitory substance type II receptor (MISIIR), sialylated Tn antigen, fibroblast activation antigen (FAP), endosialin (CD248), epidermal growth factor receptor variant III (EGFRvIII), tumor-associated antigen L6 (TAL6), CD-63, TAG-72, Thomsen-Friedenreich antigen (TF-antigen), insulin-like growth factor I receptor (IGF-IR), Cora antigen, CD7, CD22, CD79a, CD79b, G250, F19, EphA2, and MT-MM. In certain embodiments, the present invention provides chimeric polypeptide assembly compositions comprising a first portion binding domain having binding affinity for a tumor-specific marker comprising anti-marker VL and VH sequences. Illustrative, non-limiting examples of VL and VH sequences specific for certain of these tumor markers are listed in Table 6f.In another embodiment, the present invention provides a chimeric polypeptide assembly composition comprising a first portion binding domain with binding affinity to a tumor-specific marker, comprising a CDR-L1 region, a CDR-L2 region, a CDR-L3 region, a CDR-H1 region, a CDR-H2 region, and a CDR-H3 region, each derived from a respective VL and VH sequence. In a preferred embodiment, the binding domains are determined by an in vitro binding assay. -7 ~10 -10 K less than M d It has a value.
[0214] It is specifically contemplated that a chimeric polypeptide assembly composition can comprise any one of the aforementioned binding domains or sequence variants thereof, so long as the variant exhibits binding specificity for the desired antigen. In one embodiment, sequence variants are created by substituting an amino acid in the VL or VH sequence with a different amino acid. In deletion variants, one or more amino acid residues in the VL or VH sequence described herein are removed. Thus, deletion variants include all fragments of the binding domain polypeptide sequence. In substitution variants, one or more amino acid residues in the VL or VH (or CDR) polypeptide are removed and replaced with different residues. In one aspect, the substitutions are conservative in nature, and conservative substitutions of this type are well known in the art. In addition, it is specifically contemplated that a composition comprising the first and second binding domains disclosed herein can be utilized in any of the methods disclosed herein.
[0215] unstructured 3D structure Typically, the XTEN polypeptide component of the fusion proteins disclosed herein is designed to behave like a denatured peptide sequence under physiological conditions, despite the extended length of the polymer. "Denatured" describes the state of a peptide in solution, characterized by a large conformational freedom of the peptide backbone. Most peptides and proteins adopt a denatured conformation in the presence of high concentrations of denaturants or at elevated temperatures. Peptides in the denatured conformation have, for example, characteristic circular dichroism (CD) spectra, as determined by NMR. The term "denatured conformation" and "amorphous conformation" are used interchangeably herein. In some cases, the invention provides XTEN polypeptides that may resemble a predominantly deleted denatured sequence in secondary structure under physiological conditions. In other cases, the XTEN polypeptide may substantially lack secondary structure under physiological conditions. "Predominantly deleted," as used in this context, means that less than 50% of the XTEN amino acid residues of each XTEN polypeptide contribute to secondary structure as measured or determined by the means described herein. "Substantially deleted," as used in this context, means that at least about 60%, or about 70%, or about 80%, or about 90%, or about 95%, or at least about 99% of the XTEN amino acid residues of the XTEN sequence do not contribute to secondary structure as measured or determined by the means described herein.
[0216] Various methods for identifying the presence or absence of secondary and tertiary structure in a given polypeptide have been established in the art. In particular, XTEN secondary structure can be measured spectrophotometrically, for example, by circular dichroism spectroscopy in the far-ultraviolet spectral region (190-250 nm). Secondary structural elements such as alpha helices and beta sheets each give rise to characteristic shapes and sizes in CD spectra. Secondary structure can also be predicted for polypeptide sequences through certain computer programs or algorithms, such as the well-known Chou-Fasman algorithm (Chou, PY, et al. (1974) Biochemistry, 13:222-45) and the Garnier-Osguthorpe-Robson ("GOR") algorithm (Garnier J, Gibrat JF, Robson B. (1996), GOR method for predicting protein secondary structure from amino acid sequence. Methods Enzymol 266:540-553), described in U.S. Patent Publication No. 20030228309A1. For a given sequence, the algorithm can predict the presence or absence of some secondary structure, expressed, for example, as the total and / or percentage of residues in the sequence that form alpha-helices or beta-sheets, or the percentage of residues in the sequence that are predicted to result in random coil formation (lacking secondary structure).
[0217] In some cases, the XTEN polypeptides used in the fusion protein compositions can have an alpha-helix percentage ranging from 0% to less than about 5%, as determined by the Chou-Fasman algorithm. In other cases, the XTEN polypeptides comprising the fusion protein compositions can have a beta-sheet percentage ranging from 0% to less than about 5%, as determined by the Chou-Fasman algorithm. In some cases, the XTEN sequences of the fusion protein compositions can have an alpha-helix percentage ranging from 0% to less than about 5% and a beta-sheet percentage ranging from 0% to less than about 5%, as determined by the Chou-Fasman algorithm. In preferred embodiments, the XTEN polypeptides comprising the fusion protein compositions can have an alpha-helix percentage less than about 2% and a beta-sheet percentage less than about 2%. In other cases, the XTEN sequences of the fusion protein compositions can have a high degree of random coil percentage, as determined by the GOR algorithm. In some embodiments, the XTEN polypeptide can have at least about 80%, more preferably at least about 90%, more preferably at least about 91%, more preferably at least about 92%, more preferably at least about 93%, more preferably at least about 94%, more preferably at least about 95%, more preferably at least about 96%, more preferably at least about 97%, more preferably at least about 98%, and most preferably at least about 99% random coil as determined by the GOR algorithm.
[0218] Net Charge In other cases, the XTEN polypeptide incorporates amino acid residues that have a net charge. and / or may have amorphous characteristics imparted by a reduced proportion of hydrophobic amino acids in the XTEN polypeptide. The overall net charge and net charge density can be controlled by modifying the content of charged amino acids in the XTEN polypeptide. In some cases, the net charge density of the XTEN of the composition can be above +0.1 or below -0.1 charges / residue. In other cases, the net charge of the XTEN polypeptide can be about 0%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, or about 20% or more.
[0219] Because most human or animal tissues and surfaces have a net negative charge, XTEN polypeptides can be designed to have a net negative charge to minimize nonspecific interactions between XTEN polypeptide-containing compositions and various surfaces, such as blood vessels, healthy tissues, or various receptors. Without being bound by theory, XTEN polypeptides individually carry a high net negative charge and can adopt an open conformation due to electrostatic repulsion between individual amino acids of the XTEN polypeptide distributed throughout the sequence of the XTEN polypeptide. This distribution of net negative charges over the extended sequence length of the XTEN polypeptide can result in an amorphous conformation, which in turn can result in an effective increase in hydrodynamic radius. Thus, in one embodiment, the invention provides XTEN polypeptides containing about 8, 10, 15, 20, 25, or even about 30% glutamic acid. The XTEN polypeptides of the compositions of the invention generally have no or a low degree of positively charged amino acids. In some cases, an XTEN polypeptide may have less than about 10% of the amino acid residues with a positive charge, or less than about 7%, or less than about 5%, or less than about 2% of the amino acid residues with a positive charge. However, the present invention contemplates constructs in which a limited number of positively charged amino acids, such as lysine, can be incorporated into the XTEN polypeptide to permit conjugation between the epsilon amine of the lysine and a reactive group on a peptide, a linker bridge, or a reactive group on a drug or small molecule to be conjugated to the XTEN polypeptide backbone. In the foregoing, fusion proteins containing one or more XTEN polypeptides, a biologically active protein, and a chemotherapeutic agent useful in the treatment of metabolic diseases or disorders can be constructed, with the maximum number of agent molecules incorporated into the XTEN polypeptide component being determined by the number of lysines or other amino acids with reactive side chains (e.g., cysteine) incorporated into the XTEN.
[0220] In some cases, XTEN polypeptides contain charged residues separated by other residues, such as serine or glycine, which may lead to better expression or purification behavior. Based on net charge, XTEN polypeptides of human or animal compositions can have an isoelectric point (pI) of 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, or 6.5. In preferred embodiments, XTEN polypeptides have an isoelectric point between 1.5 and 4.5. In these embodiments, the XTEN incorporated into the BPXTEN fusion protein compositions of the invention will have a net negative charge under physiological conditions, which may contribute to amorphous conformation and reduced binding of the XTEN polypeptide component to mammalian proteins and tissues.
[0221] Because hydrophobic amino acids can impart structure to a polypeptide, the present invention provides that the hydrophobic amino acid content in an XTEN polypeptide is typically less than 5%, or less than 2%, or less than 1%. In one embodiment, the methionine and tryptophan amino acid content in the XTEN component of a BPXTEN fusion protein is typically less than 5%, or less than 2%, and most preferably less than 1%. In another embodiment, the XTEN polypeptide has less than 10% amino acid residues that are positively charged, or less than about 7%, or less than about 5%, or less than about 2% amino acid residues that are positively charged, and the sum of methionine and tryptophan residues is less than 2%. and the sum of asparagine and glutamine residues is less than 10% of the total XTEN polypeptide.
[0222] Increased hydrodynamic radius In some embodiments, XTEN polypeptides can have a long hydrodynamic radius, conferring a corresponding increase in apparent molecular weight to BPXTEN fusion proteins incorporating the XTEN polypeptide. Linking an XTEN polypeptide to a BP sequence can result in BPXTEN compositions that can have an increased hydrodynamic radius, increased apparent molecular weight, and increased apparent molecular weight coefficient compared to a BP that is not linked to an XTEN polypeptide. For example, in therapeutic applications where extended half-life is desirable, compositions in which an XTEN polypeptide with a large hydrodynamic radius is incorporated into a fusion protein containing one or more BPs effectively expands the hydrodynamic radius of the composition beyond the glomerular pore size of approximately 3-5 nm (corresponding to an apparent molecular weight of approximately 70 kDA) (Caliceti. 2003. Pharmacokinetic and biodistribution properties of poly(ethylene glycol)-protein conjugates. Adv. Drug Deliv. Rev. 55:1261-1277), thereby resulting in reduced renal clearance of circulating proteins. Without being bound by any particular theory, XTEN polypeptides may adopt an open structure due to the inherent flexibility conferred by specific amino acids in the sequence, which lack the potential for electrostatic repulsion between the individual charges of the peptide or the potential for secondary structure. The open, extended, and amorphous structure of XTEN polypeptides may have a larger partial hydrodynamic radius compared to polypeptides of comparable sequence length and / or molecular weight with secondary and / or tertiary structure, such as typical globular proteins. Methods for determining hydrodynamic radius, such as by using size exclusion chromatography (SEC) as described in U.S. Patent Nos. 6,406,632 and 7,294,513, are well known in the art. Increasing the length of the XTEN polypeptide results in a partial increase in the parameters of hydrodynamic radius, apparent molecular weight, and apparent molecular weight coefficient, which allows BPXTEN to be tailored to a desired characteristic cutoff apparent molecular weight or hydrodynamic radius.Thus, in certain embodiments, a BPXTEN fusion protein can be comprised of an XTEN polypeptide, and the fusion protein can have a hydrodynamic radius of at least about 5 nm, or at least about 8 nm, or at least about 10 nm, or 12 nm, or at least about 15 nm. In the foregoing embodiments, the large hydrodynamic radius comprised by the XTEN polypeptide in the BPXTEN fusion protein can lead to decreased renal clearance of the resulting fusion protein, leading to a corresponding increase in terminal half-life, increased mean residence time, and / or increased renal clearance rate.
[0223] In another embodiment, XTEN polypeptides of selected lengths and sequences can be selectively incorporated into BPXTEN to generate fusion proteins with apparent molecular weights under physiological conditions of at least about 150 kDa, or at least about 300 kDa, or at least about 400 kDa, or at least about 500 kDa, or at least about 600 kDa, or at least about 700 kDa, or at least about 800 kDa, or at least about 900 kDa, or at least about 1000 kDa, or at least about 1200 kDa, or at least about 1500 kDa, or at least about 1800 kDa, or at least about 2000 kDa, or at least about 2300 kDa or more. In another embodiment, an XTEN polypeptide of selected length and sequence can be selectively linked to a BP to result in a BPXTEN fusion protein that, under physiological conditions, has an apparent molecular weight factor of at least 3, alternatively at least 4, alternatively at least 5, alternatively at least 6, alternatively at least 8, alternatively at least 10, alternatively at least 15, or an apparent molecular weight factor of at least 20 or greater. In another embodiment, the BPXTEN fusion protein, under physiological conditions, has an apparent molecular weight factor of at least about 4 relative to the actual molecular weight of the fusion protein. has an apparent molecular weight factor of about 20, or about 6 to about 15, or about 8 to about 12, or about 9 to about 10. In some embodiments, the (fusion) polypeptide exhibits an apparent molecular weight factor of greater than about 6 under physiological conditions.
[0224] Prolonged terminal half-life In some embodiments, the (fusion) polypeptide has a terminal half-life that is at least 2-fold longer, or at least 3-fold longer, or at least 4-fold longer, or at least 5-fold longer than a biologically active polypeptide that is not linked to an XTEN polypeptide. In some embodiments, the (fusion) polypeptide has a terminal half-life that is at least 2-fold longer than a biologically active polypeptide that is not linked to an XTEN polypeptide.
[0225] Administration of a therapeutically effective dose of any of the BPXTEN fusion protein embodiments described herein to a human or animal in need thereof may result in an extension of the time spent within the therapeutic window for the fusion protein by at least two-fold, or at least three-fold, or at least four-fold, or at least five-fold or more compared to administration of the corresponding BP that is not linked to an XTEN polypeptide and at an equivalent dose to a human or animal.
[0226] Low immunogenicity In another aspect, the invention provides compositions in which the XTEN polypeptide has a low degree of immunogenicity or is substantially non-immunogenic. Several factors can contribute to the low immunogenicity of an XTEN polypeptide, such as its substantially non-repetitive sequence, its amorphous structure, its high degree of solubility, its low degree or lack of self-aggregation, its low degree or lack of proteolytic sites within the sequence, and its low degree or lack of epitopes in the XTEN polypeptide.
[0227] Those skilled in the art will generally understand that polypeptides with highly repetitive short amino acid sequences (e.g., a 200 amino acid long sequence contains an average of 20 repeats or a limited set of 3- or 4-mers) and / or with consecutive repeated amino acid residues (e.g., a 5- or 6-mer sequence has identical amino acid residues) have a tendency to aggregate or form higher order structures or form tight junctions that result in crystalline or pseudocrystalline structures.
[0228] In some embodiments, the XTEN polypeptide is substantially non-repetitive, the XTEN amino acid sequence does not have three consecutive amino acids of the same amino acid type unless the amino acid is serine, in which case no more than three consecutive amino acids may be serine residues, and the XTEN amino acid sequence does not contain a sequence of three amino acids (a 3-mer) that occurs more than 16 times, more than 14 times, more than 12 times, or more than 10 times within the 200 amino acid long sequence of the XTEN polypeptide. One of skill in the art will understand that such substantially non-repetitive sequences have little tendency to aggregate, thus enabling the design of long sequence XTEN with a relatively low frequency of charged amino acids that would be likely to aggregate if the sequence or amino acid residues were otherwise more repetitive.
[0229] Conformational epitopes are formed by regions on the protein surface that consist of multiple, non-contiguous amino acid sequences of a protein antigen. Correct protein folding organizes these sequences into well-defined, stable spatial arrangements, or epitopes, that are recognized as "foreign" by the host humoral immune system, resulting in the production of antibodies against the protein or the elicitation of a cell-mediated immune response. In the latter case, an individual's immune response to a protein is largely influenced by T-cell epitope recognition, which is a function of the peptide-binding specificity of that individual's HLA-DR allotype. The cognate T-cell receptor on the T-cell surface Engagement of MHC class II peptide complexes by receptors, together with cross-linking of certain other co-receptors such as the CD4 molecule, can induce an activated state in T cells. Activation leads to the release of cytokines, which further activate other lymphocytes, such as B cells, to produce antibodies or to activate T killer cells as a complete cellular immune response.
[0230] The ability of a peptide to bind to a given MHC class II molecule for presentation on the surface of an APC (antigen-presenting cell) depends on many factors, particularly its primary sequence. In one embodiment, a lower degree of immunogenicity can be achieved by designing an XTEN polypeptide that resists antigen processing in antigen-presenting cells and / or by selecting a sequence that does not bind sufficiently to MHC receptors. The present invention provides BPXTEN fusion proteins having substantially non-repetitive XTEN polypeptides designed to reduce binding to MHC II receptors and avoid the formation of epitopes for T cell receptor or antibody binding, thereby resulting in a lower degree of immunogenicity. Avoidance of immunogenicity is, in part, a direct result of the conformational flexibility of the XTEN polypeptide, i.e., the lack of secondary structure due to the selection and order of amino acid residues. For example, of particular interest are sequences that have a low tendency to adopt a compactly folded conformation in aqueous solution or under physiological conditions that may result in conformational epitopes. Administration of fusion proteins containing XTEN polypeptides using conventional therapeutic practices and administration generally does not result in the formation of neutralizing antibodies to the XTEN polypeptides and may also reduce the immunogenicity of the BP fusion partner in the BPXTEN composition.
[0231] In one embodiment, the XTEN polypeptide utilized in a human or animal fusion protein may be substantially free of epitopes recognized by human T cells. Removal of such epitopes to produce less immunogenic proteins has been previously disclosed, see, e.g., WO 98 / 52976, WO 02 / 079232, and WO 00 / 3317, which are incorporated herein by reference. Assays for human T cell epitopes have been described (Stickler, M., et al.). (2003) J Immunol Methods, 281:95-108). Of particular interest are peptide sequences that can oligomerize without generating T cell epitopes or nonhuman sequences. This can be achieved by testing direct repeats of these sequences for the presence of T cell epitopes and the occurrence of 6- to 15-mer, and especially nonhuman 9-mer, sequences, and then modifying the design of the XTEN polypeptide to remove or destroy the epitope sequences. In some cases, the XTEN polypeptide is substantially non-immunogenic due to the limitation of the number of epitopes in the XTEN polypeptide that are predicted to bind to MHC receptors. With a reduction in the number of epitopes capable of binding to MHC receptors, there is a concomitant reduction in T cell activation and potential T cell helper function, a reduction in B cell activation or upregulation, and a reduction in antibody production. Low-level predicted T cell epitopes can be determined by epitope prediction algorithms such as, for example, T epitopes (Sturniolo, T., et al. (1999) Nat Biotechnol, 17:555-61), shown in Example 74 of International Patent Application Publication No. WO 2010 / 144502 A2, which is incorporated by reference in its entirety. The TEPITOPE score of a given peptide frame within a protein is calculated by the K of binding of that peptide frame to multiple of the most common human MHC alleles, as disclosed in Sturniolo, T., et al. (1999) Nature Biotechnology 17:555). d The score is the logarithm of the dissociation constant, affinity, and off-rate. The score should be at least 20 logs, approximately 10 to approximately -10 (10e 10K d ~10e -10 K d The peptide's T epitope score can be reduced by avoiding hydrophobic amino acids that can serve as anchor residues during peptide presentation on the MHC, such as M, I, L, V, and F, across the full complement of amino acids (corresponding to the effective constraints of the full complement of amino acids). In some embodiments, the XTEN polypeptide incorporated into BPXTEN has a predicted T cell epitope score of about -5 or greater, or -6 or greater, or -7 or greater, or -8 or greater, or a T epitope score of -9 or greater. As used herein, a score of "-9 or less" would encompass T epitope scores of 10 to -9, inclusive, but would not encompass a score of -10, since -10 is less than -9.
[0232] In another embodiment, XTEN polypeptides of the invention, including those incorporated into human or animal BPXTEN fusion proteins, can be substantially reduced in immunogenicity by limiting known proteolytic sites from the sequence of the XTEN polypeptide, thereby reducing processing of the XTEN polypeptide into small peptides capable of binding to MHC II receptors. In another embodiment, XTEN polypeptides can be substantially reduced in immunogenicity by using a sequence that is substantially devoid of secondary structure, thereby conferring resistance to many proteases due to a high-entropy structure. Thus, reducing the T epitope score and removing known proteolytic sites from the XTEN polypeptide can render XTEN-polypeptide compositions, including the XTEN polypeptide of the BPXTEN fusion protein composition, substantially incapable of binding by mammalian receptors, including those of the immune system. In one embodiment, the XTEN polypeptide of the BPXTEN fusion protein has a binding affinity of >100 nM to mammalian receptors. K d Binding or K greater than 500 nM to mammalian cell surface or circulating polypeptide receptors d , or a K greater than 1 μM d may have:
[0233] In addition, the substantially non-repetitive sequence and corresponding lack of epitopes of such embodiments of the XTEN polypeptide may limit the ability of B cells to bind to or be activated by the XTEN polypeptide. While an XTEN polypeptide may contact many different B cells over its extended sequence, each individual B cell may only contact an individual XTEN polypeptide once or several times. As a result, XTEN polypeptides may typically have a much lower tendency to stimulate B cell proliferation and thus an immune response. In one embodiment, BPXTEN may have reduced immunogenicity compared to the corresponding unfused BP. In one embodiment, administration of up to three parenteral doses to a mammal may result in detectable anti-BPXTEN IgG at a serum dilution of 1:100, but not at a dilution of 1:1000. In another embodiment, administration of up to three parenteral doses to a mammal may result in detectable anti-BP IgG at a serum dilution of 1:100, but not at a dilution of 1:1000. In another embodiment, administration of up to three parenteral doses to a mammal may result in detectable anti-XTEN IgG at a serum dilution of 1:100, but not at a dilution of 1:1000. In the foregoing embodiment, the mammal may be a mouse, rat, rabbit, or cynomolgus monkey.
[0234] An additional feature of certain embodiments of XTEN polypeptides having substantially non-repetitive sequences, as compared to those with less non-repetitive sequences (e.g., those with three consecutive amino acids that are identical), is that the non-repetitive XTEN polypeptides form weaker contacts (e.g., monovalent interactions) with antibodies, which allows the BPXTEN composition to remain in circulation for a longer period of time, resulting in a reduced likelihood of immune clearance.
[0235] In some embodiments, the (fusion) polypeptide is less immunogenic compared to a biologically active polypeptide that is not linked to an XTEN polypeptide, where immunogenicity is confirmed by measuring the production of IgG antibodies that selectively bind to the biologically active polypeptide after administration of equivalent doses to humans or animals.
[0236] Spacer and BP release segment In some embodiments, at least a portion of the biological activity of each BP is retained by intact BPXTEN. In some embodiments, the BP component is biologically active upon its release from the XTEN polypeptide by optional cleavage of sequences incorporated within the spacer sequence into BPXTEN, as described more fully herein below. They either become active or have increased biological activity.
[0237] Optional spacer sequences are optional in fusion proteins encompassed by the present invention. Spacers may be provided to enhance expression of the fusion protein from host cells or to increase steric hindrance, which allows the BP component to assume its desired tertiary structure and / or properly interact with its target molecule. For information on spacers and methods for identifying desirable spacers, see, for example, George, et al. (2003) Protein Engineering 15:871-879, specifically incorporated herein by reference. In one embodiment, the spacer comprises one or more peptide sequences that are 1 to 50 amino acid residues in length, or about 1 to 25 residues, or about 1 to 10 residues in length. Spacer sequences that do not include a cleavage site may comprise any of the 20 naturally occurring L-amino acids, preferably hydrophilic amino acids that are not sterically hindered, including, but not limited to, glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), and proline (P). In some embodiments, the spacer may be polyglycine or polyalanine, or a mixture of predominantly glycine and alanine residues. A spacer polypeptide that does not include a cleavage sequence is predominantly substantially devoid of secondary structure. In one embodiment, one or both spacer sequences in a BPXTEN fusion protein composition may further contain a cleavage sequence, which may be the same or different, that can be acted upon by a protease to release the BP from the fusion protein.
[0238] In some cases, the incorporation of a cleavage sequence into BPXTEN is designed to allow the release of a BP that becomes active or more active upon its release from the XTEN polypeptide. The cleavage sequence is located sufficiently close to the BP sequence, generally within 18, 12, 6, or 2 amino acids of the end of the BP sequence, so that any remaining residues attached to the BP after cleavage do not significantly interact with the activity of the BP (e.g., receptor binding), allowing for sufficient evaluation of proteases that can act to cleave the cleavage sequence. In some embodiments, the cleavage site is a sequence that can be cleaved by a protease endogenous to a mammal, human, or animal, and BPXTEN can be cleaved after administration to a human or animal. In such cases, BPXTEN can serve as a prodrug or circulatory depot for the BP. Examples of cleavage sites contemplated by the present invention include, but are not limited to, polypeptide sequence cleavage by a mammalian endogenous protease selected from FXIa, FXIIa, kallikrein, FVIIa, FIXa, FXa, FIIa (thrombin), elastase-2, granzyme B, MMP-12, MMP-13, MMP-17, or MMP-20, or a non-mammalian protease such as TEV, enterokinase, PreScission™ protease (rhinovirus 3C protease), and sortase A. Sequences known to be cleaved by the aforementioned proteases are known in the art. Exemplary cleavage sequences and cleavage sites within the sequences are listed in Table 7a, as well as sequence variants. For example, thrombin (activated coagulation factor II) acts on the sequence LTPRSLLV (SEQ ID NO: 222), which is cleaved after arginine at position 4 in the sequence {Rawlings ND, et al. (2008) Nucleic Acids Res., 36:D320}. Activated FIIa is generated by cleavage of FII by FXa in the presence of phospholipids and calcium, and is located downstream from factor IX in the coagulation pathway. Its natural role in activated coagulation is to cleave fibrinogen, which then in turn initiates clot formation. FIIa activity is tightly regulated, occurring only when coagulation is necessary for proper hemostasis.However, because coagulation is an ongoing process in mammals, incorporation of the LTPRSLLV (SEQ ID NO: 223) sequence into BPXTEN between the BP and the XTEN polypeptide allows the XTEN polypeptide to be removed from the adjacent BP, which occurs simultaneously with activation of either the extrinsic or intrinsic coagulation pathway when coagulation is physiologically needed, thereby releasing the BP over time, similarly to BPXTEN, for action by endogenous proteases. Incorporation of other sequences may result in sustained release of the BP and, in certain cases, a greater degree of activity of the BP from the "prodrug" form of BPXTEN.
[0239] In some cases, only two or three amino acids flanking either side of the cleavage site (a total of four to six amino acids) will be incorporated into the cleavage sequence. In other cases, the known cleavage sequence can have one or more deletions or insertions or one, two, or three amino acid substitutions of any one, two, or three amino acids in the known sequence, where the deletions, insertions, or substitutions result in decreased or increased susceptibility to proteases, but do not result in the absence of susceptibility to proteases, thereby providing the ability to tailor the release rate of the BP from the XTEN. Exemplary substitutions are shown in Table 7a. [Table 20]
[0240] In some embodiments, the BPXTEN fusion protein can include a spacer sequence that can further include one or more cleavage sequences that, when acted upon by a protease, are formed to release the BP from the fusion protein. In some embodiments, the one or more cleavage sequences can have at least about 80% (e.g., at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100%) sequence identity to a sequence listed in Table 7a.
[0241] In some embodiments, the present disclosure provides a method for the preparation of a compound that is a substrate for one or more mammalian proteases associated with or produced by cells found in or near the diseased tissue. The present invention relates to release segment peptides (or release segments (RS)). Such proteases include, but are not limited to, classes of proteases such as metalloproteinases, cysteine proteases, aspartic acid proteases, and serine proteases, including, but not limited to, those listed in Table 7b. RSs are particularly useful for incorporation into human or animal recombinant polypeptides to yield prodrug forms that can be activated upon cleavage of the RS by a mammalian protease. As described herein, an RS is incorporated into a human or animal recombinant polypeptide composition, linking the incorporated binding moiety to an XTEN (the configuration of which is described more fully below); upon cleavage of the RS by the action of one or more proteases for which the RS is a substrate, the binding moiety and XTEN are released from the composition and binding site, are no longer shielded by the XTEN, and fully regain their ability to bind their ligand. When first and second antibody fragments are incorporated into the recombinant polypeptide composition, the composition is referred to herein as an activatable antibody composition (AAC). [Table 21-1] [Table 21-2]
[0242] In one embodiment, the present disclosure provides an activatable recombinant polypeptide comprising a first release segment (RS1) sequence that, when optimally aligned, has at least 88%, or at least 94%, or 100% sequence identity to a sequence selected from the sequences set forth in Table 8a, where RS1 is a substrate for one or more mammalian proteases. In another embodiment, the present disclosure provides an activatable recombinant polypeptide comprising an RS1 and a second release segment (RS2) sequence that, when optimally aligned, each have at least 88%, or at least 94%, or 100% sequence identity to a sequence selected from the sequences set forth in Table 8a, where RS1 and RS2 are each substrates for one or more mammalian proteases. In another embodiment, the present disclosure provides an activatable recombinant polypeptide comprising a first RS (RS1) sequence that, when optimally aligned, has at least 90%, at least 93%, at least 97%, or 100% identity to a sequence selected from the sequences set forth in Table 8b, where RS is a substrate for one or more mammalian proteases. In other embodiments, the present disclosure provides an activatable recombinant polypeptide comprising an RS1 and a second release segment (RS2) sequence that, when optimally aligned, each have at least 88%, or at least 94%, or 100% sequence identity to a sequence selected from the sequences set forth in Table 8b, wherein RS1 and RS2 are substrates for one or more mammalian proteases. In embodiments of an activatable recombinant polypeptide comprising RS1 and RS2, the two release segments can be identical or the sequences can be different.
[0243] The present disclosure provides methods for treating inflammatory bowel diseases, including metalloproteinases, cysteine-binding proteins, and proteases, including the proteases listed in Table 7b. The release segment contemplates a release segment that is a substrate for one, two, or three different types of proteases selected from phosphoproteases, aspartic acid proteases, and serine proteases. In a particular aspect, the RS serves as a substrate for a protease found in close association with or co-localized with diseased tissue or cells, such as, but not limited to, tumors, cancer cells, and inflamed tissues; upon cleavage of the RS, binding moieties that would otherwise be masked by the XTEN of the human or animal recombinant polypeptide composition (and therefore have lower binding affinity for their respective ligands) are released from the composition and fully regain their ability to bind to target and / or effector cell ligands. In another embodiment, the RS of the human or animal recombinant polypeptide composition comprises an amino acid sequence that is a substrate for a cellular protease located within the targeted cell, including, but not limited to, a protease listed in Table 7b. In another particular aspect of the human or animal recombinant polypeptide composition, an RS that is a substrate for two or three types of proteases is designed with a sequence that is capable of being cleaved at different locations in the RS sequence by different proteases. Thus, an RS that is a substrate for two, three, or more classes of proteases will have two, three, or multiple distinct cleavage sites in the RS sequence, yet cleavage by a single protease will nevertheless result in the release of the binding moiety and the XTEN from the recombinant polypeptide composition comprising the RS.
[0244] In one embodiment, the RS of the present disclosure for incorporation into a human or animal recombinant polypeptide composition is selected from the group consisting of meprin, neprilysin (CD10), PSMA, BMP-1, a disintegrin and metalloproteinase (ADAM), ADAM8, ADAM9, ADAM10, ADAM12, ADAM15, ADAM17 (TACE), ADAM19, ADAM28 (MDC-L), ADAM with thrombospondin motifs (ADAMTS), ADAMTS1, ADAMTS4, ADAMTS5, MMP-1 (collagenase 1), matrix metalloproteinase-1 (MMP-1), matrix metalloproteinase-2 (MMP-2, gelatinase A), matrix metalloproteinase-3 (MMP-3, stromelysin 1), matrix metalloproteinase-7 (MMP-7, matrilysin 1), matrix metalloproteinase-8 (MMP-8, collagenase 2), matrix metalloproteinase-9 (MMP-9, gelatinase B), matrix metalloproteinase-10 (MMP-10, stromelysin 2), matrix metalloproteinase-11 (MMP-11, stromelysin 3), matrix metalloproteinase-12 (MMP-12, macrophage elastase), matrix metalloproteinase-13 (MMP-13, collagenase 3), matrix metalloproteinase-14 (MMP-1 4, MT1-MMP), matrix metalloproteinase-15 (MMP-15, MT2-MMP), matrix metalloproteinase-19 (MMP-19), matrix metalloproteinase-23 (MMP-23, CA-MMP), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-26 (MMP-26, matrilysin 2 ...24 (MMP-24, MT5-MMP), matrix metalloproteinase-26 (MMP-26, matrilysin 2), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-26 (MMP-26, matrilysin 2), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-26 (MMP-26, matrilysin 2), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-24 (MMP-24, matrilysin 2), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-24 (MMP-24, matrilysin 2), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-24 (MMP-24, matrilysin 2), matrix metalloproteinase Methoproteinase-27 (MMP-27, CMMP), legumain, cathepsin B, cathepsin C, cathepsin K, cathepsin L, cathepsin S, cathepsin X, cathepsin D, cathepsin E, secretase, urokinase (uPA), tissue-type plasminogen activator (tPA), plasmin, thrombin, prostate-specific antigen (PSA, KLK3), human neutrophil elastase (HNE), elastase, tryptase,The RS is a substrate of one or more proteases, including type II transmembrane serine protease (TTSP), DESC1, hepsin (HPN), matriptase, matriptase-2, TMPRSS2, TMPRSS3, TMPRSS4 (CAP2), fibroblast activation protein (FAP), kallikrein-related peptidase (KLK family), KLK4, KLK5, KLK6, KLK7, KLK8, KLK10, KLK11, KLK13, and KLK14. In one embodiment, the RS is a substrate of ADAM17. In one embodiment, the RS is a substrate of BMP-1. In one embodiment, the RS is a substrate of a cathepsin. In one embodiment, In one embodiment, RS is a substrate for HtrA1. In one embodiment, RS is a substrate for legumain. In one embodiment, RS is a substrate for MMP-1. In one embodiment, RS is a substrate for MMP-2. In one embodiment, RS is a substrate for MMP-7. In one embodiment, RS is a substrate for MMP-9. In one embodiment, RS is a substrate for MMP-11. In one embodiment, RS is a substrate for MMP-14. In one embodiment, RS is a substrate for uPA. In one embodiment, RS is a substrate for matriptase. In one embodiment, RS is a substrate for MT-SP1. In one embodiment, RS is a substrate for neutrophil elastase. In one embodiment, RS is a substrate for thrombin. In one embodiment, RS is a substrate for TMPRSS3. In one embodiment, RS is a substrate for TMPRSS4. In one embodiment, the RS of the human or animal recombinant polypeptide composition is a substrate for at least two proteases: legumain, MMP-1, MMP-2, MMP-7, MMP-9, MMP-11, MMP-14, uPA, and matriptase, hi another embodiment, the RS of the human or animal recombinant polypeptide composition is a substrate for legumain, MMP-1, MMP-2, MMP-7, MMP-9, MMP-11, MMP-14, uPA, and matriptase. [Table 22-1] [Table 22-2] Table 22-3 Table 23-1 Table 23-2 Table 23-3 Table 23-4 Table 23-5 Table 23-6 Table 23-7 Table 23-8 Table 23-9 Table 23-10 Table 23-11
[0245] In another embodiment, RSs for incorporation into human or animal recombinant polypeptides can be designed to be selectively sensitive to the various proteases for which they are substrates, so that they have different cleavage rates and cleavage efficiencies. Because a given protease can be found at different concentrations in diseased tissues, including, but not limited to, tumors, blood cancers, or inflamed tissues or sites of inflammation, compared to healthy tissues or the circulation, the present disclosure provides RSs with individual amino acid sequences engineered to have higher or lower cleavage efficiencies for a given protease to ensure that the recombinant polypeptide is preferentially converted from a prodrug form to an active form (i.e., by separation and release of the binding moiety and XTEN from the recombinant polypeptide after cleavage of the RS) when in the vicinity of target cells or tissues and their co-localized proteases, compared to the cleavage rate of the RS in healthy tissues or the circulation, and the released antibody fragment-binding moiety has a higher concentration of binding to the ligand in diseased tissues compared to the prodrug form that remains in the circulation. Such selective design can improve the therapeutic index of the resulting composition and reduce side effects compared to conventional therapeutic agents that do not incorporate such site-specific activation.
[0246] As used herein, cleavage efficiency is defined as the log2 value of the ratio of the percentage of test substrate containing cleaved RS to the percentage of cleaved control substrate AC1611 when the reaction is performed with a human or animal protease enzyme in a biochemical assay (further detailed in the Examples), where the initial substrate concentration is 6 μM, the reaction is incubated at 37° C. for 2 hours and then stopped by the addition of EDTA, and the amount of digestion product and uncleaved substrate are analyzed using non-reducing SDS-PAGE to establish the ratio of the percentage cleaved. Cleavage efficiency is calculated as follows:
number
[0247] In another aspect, the disclosure provides an AAC comprising multiple RSs, wherein each RS sequence is selected from the group of sequences set forth in Table 8a, and the RSs are linked to each other by 1 to 6 amino acids selected from glycine, serine, alanine, and threonine. In one embodiment, the AAC comprises a first RS and a second RS that is different from the first RS, and each RS sequence is selected from the group of sequences set forth in Table 8a. The RSs are selected from the group of sequences set forth in Tables 8a-8b, and the RSs are linked to each other by 1-6 amino acids selected from glycine, serine, alanine, and threonine. In another embodiment, the AAC comprises a first RS, a second RS different from the first RS, and a third RS different from the first and second RSs, each sequence selected from the group of sequences set forth in Table 8a, and the first, second, and third RSs are linked to each other by 1-6 amino acids selected from glycine, serine, alanine, and threonine. It is specifically contemplated that multiple RSs of an AAC can be linked to form a sequence that can be cleaved by multiple proteases with different cleavage rates or cleavage efficiencies. In another embodiment, the present disclosure provides an AAC comprising RS1 and RS2 selected from the group of sequences set forth in Tables 8a-8b and XTEN1 and XTEN2, such as those described hereinabove or elsewhere herein, wherein RS1 is fused between XTEN1 and the binding moiety and RS2 is fused between XTEN2 and the binding moiety. It is contemplated that such compositions will be more readily cleaved by diseased target tissues that express multiple proteases compared to healthy tissues or compared to when in normal circulation, and the resulting fragments bearing the binding moiety will more readily penetrate the target tissue, e.g., tumor, and have an enhanced ability to bind and link target cells and effector cells (or, in the case of AACs designed with a single binding moiety, only target cells).
[0248] The RS of the present disclosure is useful for incorporation into recombinant polypeptides as therapeutic agents for the treatment of cancer, autoimmune diseases, inflammatory diseases, and other conditions where local activation of the recombinant polypeptide is desirable. The human or animal compositions address unmet needs and are superior in one or more aspects, including enhanced terminal half-life, targeted delivery, and improved therapeutic ratios with reduced toxicity to healthy tissues, compared to conventional antibody or bispecific antibody therapeutics that are active upon injection.
[0249] In some embodiments, the (fusion) polypeptide comprises a first release segment (RS1) positioned between the (first) XTEN and the biologically active polypeptide. In some embodiments, the polypeptide further comprises a second release segment (RS2) positioned between the biologically active polypeptide and the second XTEN. In some embodiments, RS1 and RS2 are identical in sequence. In some embodiments, RS1 and RS2 are not identical in sequence. In some embodiments, RS1 comprises an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a sequence listed in Tables 8a-8b. In some embodiments, RS2 comprises an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a sequence listed in Tables 8a-8b. In some embodiments, RS1 and RS2 are each substrates for cleavage by multiple proteases at one, two, or three cleavage sites within each release segment sequence.
[0250] reference fragment
[0251] In some embodiments, the (fusion) polypeptide further comprises one or more reference fragments releasable from the polypeptide upon digestion with a protease. In some embodiments, the one or more reference fragments each comprise a biologically active portion of the polypeptide. In some embodiments, the one or more reference fragments is a single reference fragment that differs in sequence and molecular weight from all other peptide fragments releasable from the polypeptide upon digestion of the polypeptide with a protease.
[0252] Polypeptide Mixture Disclosed herein are mixtures comprising a plurality of polypeptides of varying lengths, a mixture comprising a first set of polypeptides and a second set of polypeptides. In some embodiments, Each polypeptide in the first set of polypeptides comprises a barcode fragment (a) releasable from the polypeptide by digestion with a protease and (b) having a sequence and molecular weight that differs from the sequences and molecular weights of all other fragments releasable from the first set of polypeptides. In some embodiments, the second set of polypeptides lacks the barcode fragment of the first set of polypeptides. In some embodiments, both the first set of polypeptides and the second set of polypeptides each comprise (a) a reference fragment common to the first set of polypeptides and the second set of polypeptides and releasable by digestion with a protease. In some embodiments, the ratio of the first set of polypeptides to the polypeptides comprising the reference fragment is greater than 0.70. In some embodiments, the ratio of the first set of polypeptides to the polypeptides comprising the reference fragment is greater than 0.8, 0.9, 0.95, or 0.98. In some embodiments, the reference fragment occurs no more than once in each polypeptide in the first set of polypeptides and the second set of polypeptides. In some embodiments, the protease is a protease that cleaves C-terminal to glutamic acid residues. In some embodiments, the protease is a Glu-C protease. In some embodiments, the protease is not trypsin. In some embodiments, the polypeptides of various lengths include polypeptides comprising at least one extended recombinant polypeptide (XTEN), such as any described herein above or elsewhere herein. In some embodiments, the first set of polypeptides includes full-length polypeptides, and the barcode fragments are portions of the full-length polypeptides. In some embodiments, the full-length polypeptides are (fusion) polypeptides, such as any described herein above or elsewhere herein. In some embodiments, the barcode fragments lack (do not include) both the N-terminal and C-terminal amino acids of the full-length polypeptides. In some embodiments, the mixture of polypeptides of various lengths differs from each other due to N-terminal truncation, C-terminal truncation, or both N- and C-terminal truncation of the full-length polypeptides. In some embodiments, the first set of polypeptides and the second set of polypeptides may differ in one or more pharmacological properties.Non-limiting exemplary characteristics include:
[0253] Polypeptide Characterization Methods Disclosed herein are methods for assessing the relative abundance of a first set of polypeptides in a mixture containing polypeptides of varying lengths relative to a second set of polypeptides in the mixture, wherein (1) each polypeptide in the first set of polypeptides shares a barcode fragment that occurs once among the polypeptides, and (2) each polypeptide in the second set of polypeptides lacks the barcode fragment shared by the polypeptides in the first set, and each individual polypeptide in both the first and second set of polypeptides comprises a reference fragment. The method may include contacting the mixture with a protease to generate a plurality of proteolytic fragments resulting from cleavage of the first and second set of polypeptides, wherein the plurality of proteolytic fragments comprises a plurality of reference fragments and a plurality of barcode fragments. The method may further include determining a ratio of the amount of the barcode fragments to the amount of the reference fragments, thereby assessing the relative abundance of the first set of polypeptides relative to the second set of polypeptides. In some embodiments, a barcode fragment occurs no more than once in each polypeptide in the first set of polypeptides. In some embodiments, a reference fragment occurs no more than once in each polypeptide in the first and second set of polypeptides. In some embodiments, the plurality of proteolytic fragments comprises a plurality of reference fragments and a plurality of barcode fragments. In some embodiments, the protease cleaves the first and second sets of polypeptides (or polypeptides of various lengths) on the C-terminal side of glutamic acid residues that are not followed by proline residues. In some embodiments, the protease is a Glu-C protease. In some embodiments, the protease is not trypsin. In some embodiments, determining the ratio of the amount of the barcode fragments to the amount of the reference fragments comprises quantifying the barcode fragments and the reference fragments from the mixture of polypeptides after the mixture is contacted with the protease. In some embodiments, In some embodiments, the barcode fragment and the reference fragment are identified based on their respective masses. In some embodiments, the barcode fragment and the reference fragment are identified via mass spectrometry. In some embodiments, the barcode fragment and the reference fragment are identified via liquid chromatography-mass spectrometry (LC-MS). In some embodiments, determining the ratio of the barcode fragment to the reference fragment comprises isobaric labeling. In some embodiments, determining the ratio of the barcode fragment to the reference fragment comprises adding the mixture to one or both of the isotopically labeled reference fragment and the isotopically labeled barcode fragment. In some embodiments, the polypeptides of various lengths comprise polypeptides comprising at least one extended recombinant polypeptide (XTEN) as described herein above or elsewhere herein. In some embodiments, the XTEN is characterized in that (i) it comprises at least 150 amino acids, (ii) at least 90% of the amino acid residues of the XTEN are selected from glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), and proline (P), and (iii) it comprises at least four different types of amino acids selected from G, A, S, T, E, and P. In some embodiments, a barcode fragment, when present, is part of the XTEN. In some embodiments, the mixture of polypeptides of various lengths comprises a polypeptide described herein above or elsewhere herein. In some embodiments, the polypeptides of various lengths comprise a full-length polypeptide and truncated fragments thereof. In some embodiments, the polypeptides of various lengths consist essentially of a full-length polypeptide and truncated fragments thereof. In some embodiments, the mixture of polypeptides of various lengths differs from each other due to N-terminal truncation, C-terminal truncation, or both N- and C-terminal truncation of the full-length polypeptide. In some embodiments, the full-length polypeptide is a polypeptide described herein above or elsewhere herein. In some embodiments, the ratio of the amount of the barcode fragment to the reference fragment is greater than 0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.98, or 0.99.
[0254] Isobaric labeling-based quantification of peptides In some embodiments, isobaric labeling can be used to determine the ratio of barcode fragments to reference fragments. Those skilled in the art will recognize that isobaric labeling is a mass spectrometry strategy used in quantitative proteomics, in which peptides or proteins (or portions thereof) are labeled with various chemical groups that are isobaric (identical in mass) but vary in terms of the distribution of heavy isotopes around their structure. These tags are commonly referred to as tandem mass tags, and are designed so that during tandem mass spectrometry, the mass tags are cleaved at specific linker regions upon high-energy collision-induced dissociation (CID), thereby generating reporter ions of different masses. Those skilled in the art will recognize that one of the most common isobaric tags is an amine-reactive tag.
[0255] Enhanced ability to detect and quantitate cleavage products (e.g., via isobaric labeling) may generate knowledge than may aid in the design of manufacturing processes, including purification steps, to minimize the presence of undesired variants in the purified drug substance / product.
[0256] Recombinant production The disclosure herein includes nucleic acids. The nucleic acids can comprise polynucleotides (or polynucleotide sequences) encoding (fusion) polypeptides, such as any described herein above or elsewhere herein, or the nucleic acids can comprise the reverse complement of such polynucleotides (or polynucleotide sequences).
[0257] The present disclosure herein includes an expression vector comprising a polynucleotide sequence, such as any described in the above paragraph, and a control sequence operably linked to the polynucleotide sequence.
[0258] The present disclosure of the present invention includes a host cell comprising an expression vector as described in the paragraph above. In some embodiments, the host cell is a prokaryote. In some embodiments, the host cell is E. coli. In some embodiments, the host cell is a mammalian cell.
[0259] In another aspect, the present disclosure provides a method for producing a human or animal composition. In one embodiment, the method comprises culturing a host cell containing a nucleic acid construct encoding any of the polypeptides or XTEN-containing compositions of the embodiments described herein under conditions that promote the expression of the polypeptide or BPXTEN fusion polypeptide, and then recovering the polypeptide or BPXTEN fusion polypeptide using standard purification methods (e.g., column chromatography, HPLC, etc.), wherein the composition is recovered and at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 99% of the binding fragments of the expressed polypeptide or BPXTEN fusion polypeptide are correctly folded. In another embodiment of the production method, the expressed polypeptide or BPXTEN fusion polypeptide is recovered and at least or at least 90%, or at least 95%, or at least 97%, or at least 99% of the polypeptide or BPXTEN fusion polypeptide is recovered in a monomeric, soluble form.
[0260] In another aspect, the present disclosure relates to a method for producing polypeptides and BPXTEN fusion polypeptides at high fermentation expression levels of functional proteins using E. coli or mammalian host cells, and a method for providing expression vectors encoding constructs useful in the method for producing cytotoxically active polypeptide construct compositions at high expression levels. In one embodiment, the method includes the steps of: 1) preparing a polynucleotide encoding any of the polypeptides of the embodiments disclosed herein; 2) cloning the polynucleotide into an expression vector, which may be a plasmid or other vector under the control of transcription and translation sequences suitable for high-level protein expression in a biological system; 3) transforming a suitable host cell with the expression vector; and 4) culturing the host cell in a conventional nutrient medium under conditions suitable for expression of the polypeptide composition. If desired, the host cell is E. coli. Depending on the method, expression of the polypeptide results in a fermentation titer of expression product of at least 0.05 g / L of host cells, or at least 0.1 g / L, or at least 0.2 g / L, or at least 0.3 g / L, or at least 0.5 g / L, or at least 0.6 g / L, or at least 0.7 g / L, or at least 0.8 g / L, or at least 0.9 g / L, or at least 1 g / L, or at least 2 g / L, or at least 3 g / L, or at least 4 g / L, or at least 5 g / L, and at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 99% of the expressed protein is correctly folded. As used herein, the term "correctly folded" means that the antigen-binding fragment component of the composition has the ability to specifically bind to its target ligand.In another embodiment, the disclosure provides a method for producing a polypeptide or a BPXTEN fusion polypeptide, comprising culturing host cells comprising a vector encoding the polypeptide or a polypeptide comprising a BPXTEN fusion polypeptide under conditions effective to express the polypeptide product at a concentration greater than about 10 milligrams per gram dry weight of host cells (mg / g), or at least about 250 mg / g, or about 300 mg / g, or about 350 mg / g, or about 400 mg / g, or about 450 mg / g, or about 500 mg / g of said polypeptide in a fermentation reaction broth when the fermentation reaction reaches an optical density of at least 130 at a wavelength of 600 nm, wherein the antigen-binding fragment of the expressed protein is correctly folded. Alternatively, a method for producing a BPXTEN fusion polypeptide is provided, comprising culturing host cells comprising a vector encoding the composition under conditions effective to express the polypeptide product at a concentration of greater than about 10 milligrams per gram (mg / g) of the polypeptide in a fermentation reaction broth, or at least about 250 mg / g, or about 300 mg / g, or about 350 mg / g, or about 400 mg / g, or about 450 mg / g, or about 500 mg / g dry weight host cells when the fermentation reaction reaches an optical density of at least 130 at a wavelength of 600 nm, wherein the expressed polypeptide product is soluble.
[0261] Pharmaceutical Composition Disclosed herein is a pharmaceutical composition comprising any BPXTEN polypeptide as described herein above or elsewhere herein and one or more pharmaceutically acceptable excipients.In some embodiments, the pharmaceutical composition is formulated for intradermal, subcutaneous, intravenous, intraarterial, intraperitoneal, intraperitoneal, intravitreal, intrathecal, or intramuscular administration.In some embodiments, the pharmaceutical composition is in liquid form.In some embodiments, the pharmaceutical composition is in a device that is implanted in the eye or another body part.In some embodiments, the pharmaceutical composition is in a pre-filled syringe for single injection.In some embodiments, the pharmaceutical composition is formulated as a lyophilized powder that is reconstituted before administration.
[0262] In some embodiments, the dose is administered intradermally, subcutaneously, intravenously, intraarterially, intraperitoneally, intraperitoneally, intravitreally (or otherwise injected into the eye), intrathecally, or intramuscularly. In some embodiments, the pharmaceutical composition is administered using a device implanted in the eye or another body part. In some embodiments, the human or animal is a mouse, rat, monkey, or human.
[0263] The pharmaceutical compositions may be administered for therapy by any suitable route, and in addition, the pharmaceutical compositions may also contain other pharmaceutically active compound or compounds of the present invention.
[0264] In some embodiments, the pharmaceutical composition can be administered at a therapeutically effective dose, which, in certain cases, results in an increase in the time spent within the therapeutic window of the fusion protein compared to the corresponding BP of the fusion protein not linked to the fusion protein and administered at an equivalent dose in humans or animals.
[0265] In another embodiment, the present invention provides a method of treating a disease, disorder, or condition comprising administering the pharmaceutical composition described above to a human or animal using multiple consecutive doses of the pharmaceutical composition administered using a therapeutically effective dose regimen.
[0266] The BPXTEN polypeptides of the present invention can be formulated according to known methods for preparing pharmaceutically useful compositions, whereby the polypeptides are mixed with a pharmaceutically acceptable carrier vehicle, such as an aqueous solution or buffer, in a pharmaceutically acceptable suspension or emulsion. Therapeutic formulations are prepared for storage by mixing the active ingredient having the desired purity, as described in Remington's Pharmaceutical Sciences 16th edition, Osol, A. Ed. (1980), with physiologically acceptable carriers, excipients, or stabilizers of choice.
[0267] Medicine Kit In another aspect, the invention provides kits for facilitating the use of BPXTEN polypeptides. In one embodiment, the kit contains, in at least a first container, (a) an amount of a BPXTEN fusion protein composition sufficient to treat a disease, condition, or disorder upon administration to a human or animal in need thereof, and (b) an amount of a pharmaceutically acceptable carrier, in a formulation ready for injection or reconstitution with sterile water, buffer, or dextrose. together with labeling and handling instructions identifying the BPXTEN drug, and a sheet of the drug's approved indications, instructions for reconstituting and / or administering the BPXTEN drug for use in the prevention and / or treatment of the approved indications, appropriate dosage and safety information, and information identifying the drug's lot and expiration date. In another embodiment of the foregoing, the kit can include a second container that can have a suitable diluent for the BPXTEN composition, which provides the user with the appropriate concentration of BPXTEN to be delivered to a human or animal.
[0268] Treatment method Disclosed herein is the use of a polypeptide, such as any of those described herein above or elsewhere herein, in the preparation of a medicament for treating a disease in a human or animal. In some embodiments, the particular disease to be treated depends on the selection of the biologically active protein. In some embodiments, the disease is cancer.
[0269] Disclosed herein are methods for treating a disease in a human or animal, comprising administering to a human or animal in need thereof one or more therapeutically effective doses of a pharmaceutical composition, such as any of those described herein above or elsewhere herein. In some embodiments, the disease is cancer. In some embodiments, the pharmaceutical composition is administered to the human or animal as one or more therapeutically effective doses administered according to a dosing regimen. In some embodiments, the human or animal is a mouse, rat, monkey, or human.
[0270] The following are examples of compositions and composition evaluations of the present disclosure. It will be understood that various other embodiments can be practiced in light of the summary provided above. [Example]
[0271] Example 1. Design of barcoded XTEN with minimal mutations from generic XTEN This example describes an exemplary design approach for barcoded XTEN polypeptides by making minimal mutations to the amino acid sequence of a generic XTEN polypeptide (such as one in Table 3b herein above). Relevant criteria for making minimal mutations include one or more of the following: (a) minimizing sequence changes in the corresponding XTEN polypeptide, (b) minimizing amino acid composition changes in the corresponding XTEN polypeptide, (c) substantially maintaining the net charge in the corresponding XTEN polypeptide, (d) substantially maintaining the low immunogenicity of the corresponding XTEN polypeptide, and (e) substantially maintaining the pharmacokinetic properties provided by the XTEN polypeptide.
[0272] For example, barcoded XTENs were constructed by making one or more mutations to the generic XTENs in Table 9, including deletion of a glutamic acid residue, insertion of a glutamic acid residue, substitution of a glutamic acid residue, or substitution with a glutamic acid residue, or any combination thereof. [Table 24]
[0273] Example 2. Sequence analysis of barcoded XTEN polypeptides and their selection for fusion to biologically active polypeptides ("BPs"). This example describes the design and selection of barcoded XTEN polypeptides (and their assembly into sets of more than one barcoded XTEN) for fusion to a biologically active polypeptide. Depending on the location of the barcode fragment within the XTEN and the manner in which the XTEN polypeptide is fused to a biologically active protein to form an XTEN polypeptide-containing construct (e.g., Xtenylated protease-activated T cell engager (XPAT)), the barcode fragment can indicate cleavage of the XTEN polypeptide.
[0274] In silico GluC digestion analysis was performed on two exemplary XTEN polypeptides (XTEN864 and XTEN288_1) to quantify peptide fragments releasable upon complete GluC digestion of the XTEN polypeptide. The in silico analysis takes into account that for XTEN polypeptides with consecutive glutamic acid residues (e.g., "EE"), GluC can cleave after any one of the glutamic acid residues. As shown in the results summarized in Table 10 below, the 10-mer peptide sequence "TPGTSTEPSE (SEQ ID NO: 8880)" and the 14-mer peptide sequence "GSAPGSEPATSGSE (SEQ ID NO: 8881)" occur only once and once, respectively, in the longer XTEN864, while occurring only once in all other peptides. The peptide sequence occurs more than once in XTEN864. The 14-mer peptide sequence "GSAPGSEPATSGSE (SEQ ID NO: 8881)" also occurs once and only once in the shorter XTEN288_1.
[0275] The uniqueness of the candidate barcode is evaluated with respect to all other peptide fragments releasable from the XTEN polypeptide-containing construct. Thus, a barcode sequence in one XTEN polypeptide cannot occur anywhere else in the XTEN polypeptide-containing construct, including any other XTEN polypeptides contained therein, any biologically active proteins contained therein, or any linkages between its adjacent components. For example, Table 11 shows a table of peptide "uniqueness" for a set of two XTEN polypeptides. Due to its presence in both XXTEN864 and XTEN288, the 14-mer peptide sequence "GSAPGSEPATSGSE (SEQ ID NO: 8881)" is not unique to the set of XTEN polypeptides containing both XTEN864 and XTEN288 and therefore cannot be used as a barcode to detect cleavage in a polypeptide product containing both XTEN polypeptides.
[0276] Selecting a barcode (or set of barcodes) can further include identifying and determining the correct location or position of the candidate barcode within the XTEN polypeptide. The location or position of the candidate barcode can be associated with pharmacologically relevant information of the XTEN polypeptide (and the XTEN polypeptide-containing construct as a whole), such as truncation of the XTEN polypeptide beyond a critical length and / or deletions in the XTEN polypeptide. If XXTEN864 is placed at the N-terminus of the XTEN polypeptide-containing product and truncation of 238 amino acids from the N-terminus of the product does not significantly affect the pharmacological properties of the product, the 10-mer peptide "TPGTSTEPSE (SEQ ID NO: 8880)" can serve as a suitable barcode fragment. [Table 25] [Table 26]
[0277] Exemplary barcode peptide sequences are shown below in Table 12. These barcode sequences have the structural formula (I): AAA-Glu-barcode peptide-BBB, (where "AAA" represents Gly, Ala, Ser, Thr, or Pro, and "BBB" represents Gly, Ala, Ser, or Thr configured to facilitate efficient release of the barcode peptide by GluC digestion. Notably, insertion of each barcode peptide in XTEN can result in additional unique sequence immediately before or after the inserted barcode peptide. [Table 27]
[0278] Example 3: Design and selection of XTEN in full sequence XTENized polypeptide constructs This example describes the design of a full sequence polypeptide construct containing two XTEN polypeptides, one N-terminal and the other C-terminal.
[0279] Table 13 below shows the XTEN polypeptides used in a representative barcoded BPXTEN (containing barcoded XTEN polypeptides at both the N- and C-termini) and a reference BPXTEN (containing generic XTEN at both the N- and C-termini). In a representative barcoded BPXTEN, a barcoded XTEN polypeptide (SEQ ID NO: 8014) is fused at the N-terminus of the BP, and another barcoded XTEN polypeptide (SEQ ID NO: 8015) is fused at the C-terminus of the BP. In a reference BPXTEN, a "Ref-N" XTEN polypeptide (SEQ ID NO: 8896) is fused at the N-terminus of the BP, and a "Ref-C" XTEN polypeptide (SEQ ID NO: 8897) is fused at the C-terminus of the BP. The "Ref-N" XTEN polypeptide (SEQ ID NO: 8896) is equivalent in length to the barcoded XTEN polypeptide SEQ ID NO: 8014, and the "Ref-C" XTEN polypeptide (SEQ ID NO: 8897) is equivalent in length to the barcoded XTEN polypeptide SEQ ID NO: 8015. The barcoded BPXTEN and the reference BPXTEN each contain a reference sequence in the BP component. The reference sequence is unique and has a molecular weight different from all other peptide fragments that can be released from the corresponding BPXTEN upon complete digestion with GluC protease (e.g., as described in Example 5). The uniqueness of the reference sequence is evaluated with respect to all other peptide fragments that can be released from the BPXTEN construct. [Table 28]
[0280] Example 4: Recombinant construction and production of barcoded Xtenylated fusion polypeptides Examples 4a-4b describe the recombinant construction, production, and purification of full-length polypeptides containing barcoded XTEN polypeptides using the methods disclosed herein.
[0281] Example 4a. Xtenylated fusion polypeptides containing barcoded XTEN at the C-terminus Expression: A construct encoding an Xtenylated fusion polypeptide containing an anti-EpCAM single-chain variable fragment (scFv) at the C-terminus and an 864-amino acid barcoded XTEN sequence (SEQ ID NO: 8008) was expressed in the proprietary E. coli AmE098 strain and peripherally partitioned via an N-terminal secretory leader sequence (MKKNIAFLLASMFVFSIATNAYA-) (SEQ ID NO: 8898), which is cleaved during translocation. Fermentation cultures were grown at 37°C using animal-free complex medium, and the temperature was shifted to 26°C prior to phosphate depletion. During harvest, the fermentation whole broth was centrifuged to pellet the cells. At harvest, the total volume and wet cell weight (WCW, the ratio of pellet to supernatant) were recorded, and the pelleted cells were collected and frozen at -80°C.
[0282] Harvesting: Frozen cell pellets were lysed in lysis buffer (17%) at a target of 30% wet cell weight. The cells are resuspended in 7 mM citric acid, 22.3 mM NaHPO, 75 mM NaCl, 2 mM EDTA, pH 4.0. The resuspension is equilibrated to pH 4 and then homogenized via two passes at 800 ± 50 bar while the output temperature is monitored and maintained at 15 ± 5°C. The pH of the homogenate is confirmed to be within the specified range (pH 4.0 ± 0.2).
[0283] Clarification: To reduce endotoxins and host cell impurities, the homogenate is cold-flocculated (10±5°C) and acidified (pH 4.0±0.2) overnight (15-20 hours). To remove the insoluble fraction, the flocculated homogenate is centrifuged at 16,900 RCF for 40 minutes at 2-8°C, and the supernatant is pooled. The supernatant is diluted approximately three-fold with Milli-Q water (MQ) and then adjusted to 7±1 mS / cm with 5M NaCl. To remove nucleic acids, lipids, and endotoxins and act as a filter aid, the supernatant is adjusted to 0.1% (m / m) diatomaceous earth. To maintain the filter aid in suspension, the supernatant is mixed via impeller and allowed to equilibrate for 30 minutes. A filter train consisting of a deep filter followed by a 0.22 μm filter is assembled and then rinsed with MQ. The supernatant is pumped through the filter train while adjusting the flow rate to maintain a pressure drop of 25 ± 5 psig. To adjust the complex buffer system (based on the ratio of citric acid to NaHPO) to the desired range for capture chromatography, the filtrate is adjusted with 500 mM NaHPO for a final ratio of NaHPO to citric acid of 9.33:1, ensuring that the pH of the buffered filtrate is within the specified range (pH 7.0 ± 0.2).
[0284] purification
[0285] AEX Capture: To separate dimers, aggregates, and large truncated fragments from the monomeric product and to remove endotoxins and nucleic acids, anion exchange (AEX) chromatography is utilized to capture the negatively charged C-terminal XTEN domain. AEX1 stationary phase (GE Q Sepharose FF), AEX1 mobile phase A (12.2 mM NaHPO, 7.8 mM NaHPO, 40 mM NaCl), and AEX1 mobile phase B (12.2 mM The column is equilibrated with AEX1 mobile phase A. Based on the total protein concentration measured by bicinchoninic acid (BCA) assay, the filtrate is loaded onto a column targeting 28 ± 4 g / L of resin, chased with AEX1 mobile phase A, and then washed in steps up to 30% B. The bound material is eluted with a gradient from 30% B to 60% B over 20 CV. Fractions are collected in 1 CV aliquots with an A220 ≥ 100 mAU above baseline (local). The eluted fractions are analyzed by SDS-PAGE and SE-HPLC and pooled.
[0286] IMAC intermediate purification: To ensure C-terminal integrity, immobilized metal affinity chromatography (IMAC) is used to capture the C-terminal polyhistidine tag (His(6) (SEQ ID NO: 8031)). The following IMAC stationary phase (GE IMAC Sepharose FF), IMAC mobile phase A (18.3 mM NaHPO, 1.7 mM NaHPO, 500 mM NaCl, 1 mM imidazole), and IMAC mobile phase B (18.3 mM NaHPO, 1.7 mM NaHPO, 500 mM NaCl, 500 mM imidazole) are used here. The column is charged with zinc solution and equilibrated with IMAC mobile phase A. The AEX1 pool is adjusted to pH 7.8±0.1, 50±5 mS / cm (with 5M NaCl), and 1 mM imidazole and loaded onto an IMAC column targeting 2 g / L resin and chased with IMAC mobile phase A until the absorbance at 280 nm (A280) returns to (local) baseline. Bound material is eluted in steps up to 25% IMAC mobile phase B. IMAC elution collection begins when A280 ≥ 10 mAU (local) above baseline is pumped into a vessel preloaded with enough EDTA to bring 2 CV to 2 mM EDTA, and terminated when 2 CV have been collected. The eluate is analyzed by SDS-PAGE.
[0287] Protein-L intermediate purification: To ensure N-terminal integrity, Protein L is used to capture the kappa domain located near the N-terminus of the BPXTEN molecule (specifically, aEpCAM scFv). Protein-L stationary phase (GE Capto L) is used, along with Protein L mobile phase A (16.0 mM citric acid, 20.0 mM NaHPO, pH 4.0 ± 0.1), Protein L mobile phase B (29.0 mM citric acid, 7.0 mM NaHPO, pH 2.60 ± 0.02), and Protein L mobile phase C (3.5 mM citric acid, 32.5 mM NaHPO, 250 mM NaCl, pH 7.0 ± 0.1). The column is equilibrated with Protein L mobile phase C. The IMAC eluate is adjusted to pH 7.0±0.1 and 30±3 mS / cm (with 5 M NaCl and MQ) and loaded onto a Protein L column targeting 2 g / L resin, then chased with Protein L mobile phase C until the absorbance at 280 nm (A280) returns to the (local) baseline. The column is washed with Protein L mobile phase A, followed by a low-pH elution using Protein L mobile phases A and B. Bound material is eluted at approximately pH 3.0 and collected in a container pre-charged with 4 parts by weight of 0.5 M NaHPO for every 10 parts by weight of collected volume. Fractions are analyzed by SDS-PAGE.
[0288] HIC Polishing: To separate N-terminal variants (the four absolute N-terminal residues are not essential for Protein L binding) and overall structural variants, hydrophobic interaction chromatography (HIC) is used. The HIC stationary phase (GE Capto Phenyl ImpRes) is used, along with HIC mobile phase A (20 mM histidine, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.1) and HIC mobile phase B (1 M ammonium sulfate, 20 mM histidine, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.1). The column is equilibrated with HIC mobile phase B. The adjusted Protein L eluate is loaded onto a HIC column targeting 2 g / L resin and chased with HIC mobile phase B until the absorbance at 280 nm (A280) returns to the (local) baseline. The column is washed with 50% B. Bound material is eluted with a gradient of 50% B to 0% B over 75 CV. Fractions are collected in 1 CV aliquots with an A280 ≥ 3 mAU above baseline (local). Elution fractions are analyzed by SE-HPLC and HI-HPLC and pooled.
[0289] Formulation: The product is exchanged into formulation buffer and anion exchange is again used to capture the C-terminal XTEN to bring the product to the target concentration (0.5 g / L). AEX2 stationary phase (GE Q Sepharose FF), AEX2 mobile phase A (20 mM histidine, 40 mM NaCl, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.2), AEX2 mobile phase B (20 mM histidine, 1 M NaCl, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.2), and AEX2 mobile phase C (12.2 mM NaHPO, 7.8 mM NaHPO, 40 mM NaCl, 0.02% (w / v) polysorbate 80, pH 7.0 ± 0.2) are used here. The column is equilibrated with AEX2 mobile phase C. The HIC pool is adjusted to pH 7.0±0.1 and 7±1 mS / cm (in MQ) and loaded onto an AEX2 column targeting 2 g / L resin, then chased with AEX2 mobile phase C until A280 returns to (local) baseline. The column is washed with AEX2 mobile phase A (20 mM histidine, 40 mM NaCl, 0.02% (w / v) polysorbate 80, pH 6.5±0.2). AEX2 mobile phases A and B are used to generate and elute {NaCl} steps. Bound material is eluted in steps up to 38% AEX2 mobile phase B. AEX2 elution collection begins when A280 ≥ 5 mAU above baseline (local) and ends when two column volumes have been collected. The AEX2 eluate is 0.22 μm filtered in a BSC, aliquoted, labeled, and stored at −80°C as bulk drug substance (BDS). The bulk drug substance (BDS) is verified to meet all lot release criteria by various analytical methods. Overall quality is analyzed by SDS-PAGE, the ratio of monomer to dimer is analyzed by SE-HPLC, and N-terminal quality and product uniformity are analyzed by HI-HPLC. Further analysis.
[0290] Example 4. Xtenylated fusion polypeptides containing barcoded XTEN at the C-terminus and barcoded XTEN at the N-terminus Expression: A construct encoding an Xtenylated fusion polypeptide containing an anti-EGFR single-chain variable fragment (scFv) at the C-terminus, an 864-amino acid barcoded XTEN (SEQ ID NO: 8008), and an N-terminus barcoded XTEN (SEQ ID NO: 8007) was expressed in the proprietary E. coli AmE098 strain and peripherally partitioned via an N-terminal secretory leader sequence (MKKNIAFLLASMFVFSIATNAYA-) (SEQ ID NO: 8898), which is cleaved during translocation. Fermentation cultures were grown at 37°C using animal-free complex medium, and the temperature was shifted to 26°C prior to phosphate depletion. During harvest, the fermentation whole broth was centrifuged to pellet the cells. At harvest, the total volume and wet cell weight (WCW, the ratio of pellet to supernatant) were recorded, and the pelleted cells were collected and frozen at -80°C.
[0291] Harvesting: The frozen cell pellet is resuspended in lysis buffer (100 mM citric acid) targeting 30% wet cell weight. The resuspension is equilibrated to pH 4.4 and then homogenized at 17,000 ± 200 bar while the output temperature is monitored and maintained at 15 ± 5°C. The pH of the homogenate is confirmed to be within the specified range (pH 4.4 ± 0.1).
[0292] Clarification: To reduce endotoxins and host cell impurities, the homogenate is flocculated overnight (15-20 hours) at low temperature (10±5°C), acidic (pH 4.4±0.1). To remove the insoluble fraction, the flocculated homogenate is centrifuged at 8,000 RCF and 2-8°C for 40 minutes, and the supernatant is stored. To remove nucleic acids, lipids, and endotoxins and act as a filter aid, the supernatant is adjusted to 0.1% (m / m) diatomaceous earth. To maintain the filter aid in suspension, the supernatant is mixed via impeller and allowed to equilibrate for 30 minutes. A filter train consisting of a depth filter followed by a 0.22 μm filter is assembled and then flushed with MQ. The supernatant is pumped through the...
Claims
1. A polypeptide comprising an extended recombinant polypeptide (XTEN), (a) (i) a set of non-overlapping sequence motifs, wherein each non-overlapping sequence motif of the set is repeated at least twice in the XTEN polypeptide; and (ii) an extended recombinant polypeptide (XTEN) comprising an additional non-overlapping sequence motif that occurs only once within said XTEN polypeptide; and (b) a first barcode fragment releasable from the polypeptide upon partial or complete digestion with a protease, wherein the first barcode fragment is a portion of the XTEN polypeptide that contains the sequence motif that occurs only once within the XTEN polypeptide and that differs in sequence and molecular weight from all other peptide fragments releasable from the polypeptide upon complete digestion of the polypeptide with the protease; A polypeptide wherein the barcode fragment does not include the N-terminal amino acid or the C-terminal amino acid of the polypeptide.
2. 2. The polypeptide of claim 1, wherein the barcode fragment does not contain a glutamic acid immediately adjacent to another glutamic acid in the XTEN polypeptide.
3. The polypeptide according to any one of claims 1 to 2, wherein the barcode fragment has a glutamic acid at its C-terminus.
4. The polypeptide of any one of claims 1 to 3, wherein the barcode fragment has an N-terminal amino acid immediately following a glutamic acid residue.
5. 5. The polypeptide of claim 4, wherein the glutamic acid residue preceding the N-terminal amino acid is not immediately adjacent to another glutamic acid residue.
6. The polypeptide of any one of claims 1 to 4, wherein the barcode fragment does not contain a glutamic acid residue at a position other than the C-terminus of the barcode fragment, unless the glutamic acid is immediately followed by a proline.
7. The polypeptide of claims 1 to 6, wherein the barcode fragment is located 10 to 150 amino acids from the N-terminus or C-terminus of the polypeptide.
8. 8. The polypeptide of any one of claims 1 to 7, wherein the sequence motifs of the set of non-overlapping sequence motifs are identified by SEQ ID NOs: 182-203 and 1715-1722.
9. 9. The polypeptide of claim 8, wherein the sequence motifs of the set of non-overlapping sequence motifs are identified by SEQ ID NOs: 186-189.
10. 10. The polypeptide of claim 9, wherein the set of non-overlapping sequence motifs comprises at least two, at least three, or all four of the sequence motifs identified by SEQ ID NOs: 186-189.
11. A polypeptide comprising an extended recombinant polypeptide (XTEN), wherein the XTEN polypeptide is a portion of the XTEN polypeptide, releasable from the polypeptide upon digestion with a protease, and the sequence and molecular weight of the XTEN polypeptide are such that upon complete digestion of the polypeptide by the protease, a first barcode fragment that is distinct from all other peptide fragments releasable from the polypeptide, said barcode fragment comprising: (i) does not include the N-terminal amino acid or the C-terminal amino acid of the polypeptide; (ii) does not contain a glutamic acid immediately adjacent to another glutamic acid in said XTEN polypeptide; (iii) has a glutamic acid at its C-terminus; (iv) having an N-terminal amino acid immediately preceded by a glutamic acid residue; (v) A polypeptide located 10 to 125 amino acids from either the N-terminus or C-terminus of the polypeptide.
12. 12. The polypeptide of claim 11, wherein the glutamic acid residue preceding the N-terminal amino acid is not immediately adjacent to another glutamic acid residue.
13. 13. The polypeptide of claim 11 or claim 12, wherein the barcode fragment does not contain a glutamic acid residue at any position other than the C-terminus of the barcode fragment, unless the glutamic acid is immediately followed by a proline.
14. 12. The polypeptide of claim 11, wherein the XTEN polypeptide comprises multiple non-overlapping sequence motifs, each of the sequence motifs being 9 to 14 amino acids in length.
15. 13. The polypeptide of claim 12, wherein the sequence motifs of the set of non-overlapping sequence motifs are identified by SEQ ID NOs: 182-203 and 1715-1722.
16. 16. The polypeptide of claim 15, wherein the sequence motifs of the set of non-overlapping sequence motifs are identified by SEQ ID NOs: 186-189.
17. 17. The polypeptide of claim 16, wherein the set of non-overlapping sequence motifs comprises at least two, at least three, or all four of the sequence motifs SEQ ID NOs: 186-189.
18. 18. The polypeptide of any one of claims 1-17, wherein at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid residues of the XTEN polypeptide are glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), or proline (P).
19. 18. The polypeptide of any one of claims 1-17, wherein the XTEN polypeptide is 150-3000 amino acids in length.
20. 20. The polypeptide of claim 19, wherein the XTEN polypeptide is 150 to 1000 amino acids in length.
21. 21. The polypeptide of any one of claims 1 to 20, wherein the barcode fragment is located within 200, 150, 100, or 50 amino acids of the N-terminus of the polypeptide.
22. 22. The polypeptide of any one of claims 1 to 21, wherein the barcode fragment is located between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the N-terminus of the protein.
23. The barcode fragment is 200, 150, 100, or The polypeptide according to any one of claims 1 to 22, wherein the polypeptide is located within 50 amino acids of the amino acid sequence of the first amino acid sequence.
24. 24. The polypeptide of claim 23, wherein the barcode fragment is located between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the C-terminus of the protein.
25. The polypeptide of any one of claims 1 to 24, wherein the barcode fragment is at least 4 amino acids in length.
26. 25. The polypeptide of any one of claims 1 to 24, wherein the barcode fragment is 4 to 20, 5 to 15, 6 to 12, or 7 to 10 amino acids in length.
27. 27. The polypeptide of any one of claims 1 to 26, wherein the barcode fragment is identified by SEQ ID NOs: 8020 to 8030 (BAR001 to BAR011).
28. 28. The polypeptide of any one of claims 1-27, wherein the polypeptide further comprises a second barcode fragment, wherein the second barcode fragment is a portion of the XTEN polypeptide that contains the sequence motif that occurs only once within the XTEN polypeptide and that differs in sequence and molecular weight from all other peptide fragments that can be liberated from the polypeptide upon complete digestion of the polypeptide by the protease.
29. 29. The polypeptide of Claim 28, wherein said polypeptide further comprises a third barcode fragment, said third barcode fragment being part of said XTEN polypeptide and differing in sequence and molecular weight from all other peptide fragments liberable from said polypeptide upon complete digestion of said polypeptide by said protease.
30. 30. The polypeptide of any one of claims 1-29, wherein the XTEN polypeptide has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to a sequence identified by SEQ ID NOs:8001-8019.
31. 32. The polypeptide of any one of claims 1-31, wherein the XTEN polypeptide is at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids in length.
32. A biologically active polypeptide linked to a polypeptide according to claims 1 to 31.
33. 33. The biologically active polypeptide of claim 32, wherein the XTEN polypeptide has a proximal end and a distal end with respect to the biologically active polypeptide, the proximal end is located closer to the biologically active polypeptide than the distal end, and the barcode fragment is located within a region of the XTEN polypeptide that extends from 5% to 50%, 7% to 40%, or 10% to 30% of the length of the XTEN polypeptide, as measured from the distal end.
34. A biologically active polypeptide described in either claim 32 or claim 33, further comprising one or more reference fragments releasable from the polypeptide upon digestion with the protease, each of the one or more reference fragments comprising a portion of the biologically active polypeptide.
35. 35. The biologically active polypeptide of claim 34, wherein the one or more reference fragments is a single reference fragment that differs in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide upon digestion of the polypeptide by the protease.
36. 36. The biologically active polypeptide of any one of claims 32-35, further comprising a first release segment (RS1) located between the XTEN polypeptide and the biologically active polypeptide.
37. 37. The biologically active polypeptide of claim 36, wherein the RS1 comprises an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to a sequence identified in Tables 8a-8b.
38. 38. The biologically active polypeptide of any one of claims 32 to 37, wherein said biologically active polypeptide is identified herein by any one or combination of the sequences of Tables 4a to 4h and 6a to 6f.
39. 39. The biologically active polypeptide of any one of claims 32-38, wherein the polypeptide has a terminal half-life that is at least two-fold longer compared to the biologically active polypeptide that is not linked to any XTEN polypeptide.
40. 40. The biologically active polypeptide of any one of claims 32-39, wherein the polypeptide is less immunogenic compared to the biologically active polypeptide not linked to any XTEN polypeptide, where immunogenicity is confirmed by measuring the production of IgG antibodies that selectively bind to the biologically active polypeptide after administration of an equivalent dose to a human or animal.
41. 41. The biologically active polypeptide of any one of claims 32 to 40, wherein the polypeptide exhibits an apparent molecular weight coefficient of greater than about 6 under physiological conditions.
42. 42. The biologically active polypeptide of any one of claims 32-41, further comprising a second XTEN polypeptide, wherein the first XTEN polypeptide is located at the N-terminus of the biologically active polypeptide and the second XTEN polypeptide is located at the C-terminus of the biologically active polypeptide.
43. 43. The biologically active polypeptide of claim 42, further comprising a second release segment (RS2) located between the biologically active polypeptide and the second XTEN.
44. 44. The biologically active polypeptide of claim 43, wherein RS1 and RS2 are identical in sequence.
45. 44. The biologically active polypeptide of claim 43, wherein RS1 and RS2 are not identical in sequence.
46. 46. A biologically active polypeptide according to any one of claims 43 to 45, wherein each of RS1 and RS2 is a substrate for cleavage by multiple proteases at one, two or three cleavage sites within each release segment sequence.
47. 47. The method of claim 42, wherein the polypeptide further comprises a barcode fragment that is part of the second XTEN and that differs in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide upon complete digestion of the polypeptide by the protease. A biologically active polypeptide according to any one of claims 1 to 4.
48. 48. The biologically active polypeptide of claim 47, wherein the additional barcode fragment does not include the C-terminal amino acid of the polypeptide.
49. 49. The biologically active polypeptide of claim 47 or claim 48, wherein the further barcode fragment comprises a glutamic acid residue at its C-terminus.
50. 50. The biologically active polypeptide of any one of claims 47-49, wherein the additional barcode fragment of the second XTEN is located within 200, 150, 100, or 50 amino acids of the C-terminus of the polypeptide.
51. 51. The biologically active polypeptide of any one of claims 47-50, wherein the additional barcode fragment of the second XTEN is located between 10-200, 30-200, 40-150, or 50-100 amino acids from the C-terminus of the polypeptide.
52. 52. The biologically active polypeptide of any one of claims 47 to 51, wherein the further barcode fragment is 4 to 20, 5 to 15, 6 to 12, or 7 to 10 amino acids in length.
53. 53. The biologically active polypeptide of any one of claims 47 to 52, wherein the further barcode fragments are identified by SEQ ID NOs: 8020 to 8030 (BAR001 to BAR011).
54. 54. The biologically active polypeptide of any one of claims 47-53, further comprising a set of barcode fragments comprising the further barcode fragment and at least one additional barcode fragment, wherein each barcode fragment of the set of barcode fragments is part of the second XTEN polypeptide and differs in sequence and molecular weight from all other peptide fragments liberable from the polypeptide upon complete digestion of the polypeptide by the protease.
55. 56. The biologically active polypeptide of any one of claims 42-55, wherein the second XTEN is identified by SEQ ID NOs: 8001-8019.
56. 56. The biologically active polypeptide of any one of claims 47 to 55, wherein the further barcode fragment does not contain a glutamic acid residue immediately adjacent to another glutamic acid residue in the polypeptide.
57. 57. The biologically active polypeptide of any one of claims 42-56, wherein at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid residues of the second XTEN polypeptide are glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), or proline (P).
58. 58. The biologically active polypeptide of any one of claims 42-57, wherein the sum of the total number of amino acids in the first XTEN polypeptide and the total number of amino acids in the second XTEN polypeptide is at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, or at least 800 amino acids.
59. 59. The biologically active polypeptide of any one of claims 42-58, wherein the second XTEN polypeptide comprises multiple non-overlapping sequence motifs, and each sequence motif in the second XTEN polypeptide is 9 to 14 amino acids in length.
60. 60. The biologically active polypeptide of claim 59, wherein for the second XTEN polypeptide, the sequence motifs of the plurality of non-overlapping sequence motifs are identified by SEQ ID NOs: 182-203 and 1715-1722.
61. 61. The biologically active polypeptide of claim 60, wherein the sequence motifs of the plurality of non-overlapping sequence motifs are identified by SEQ ID NOs: 186-189.
62. 61. The biologically active polypeptide of claim 59 or claim 60, wherein for the second XTEN polypeptide, the multiple non-overlapping sequence motifs comprise at least two, at least three, or all four of the following motifs: SEQ ID NOs: 186-189.
63. 63. The biologically active polypeptide of any one of claims 42-62, wherein the second XTEN polypeptide is 150-3000 amino acids in length.
64. 64. The biologically active polypeptide of claim 63, wherein the second XTEN polypeptide is 150 to 1000 amino acids in length.
65. 65. The biologically active polypeptide of any one of claims 42-64, wherein the second XTEN polypeptide has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to a sequence identified by SEQ ID NOs:8001-8019.
66. 65. The biologically active polypeptide of any one of claims 42-64, wherein the second XTEN polypeptide is at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids in length.
67. A mixture comprising a plurality of polypeptides of various lengths, a first set of polypeptides, each polypeptide of the first set of polypeptides comprising a barcode fragment releasable from the polypeptide by digestion with a protease and having a sequence and molecular weight that differs from the sequences and molecular weights of all other fragments releasable from the first set of polypeptides; and a second set of polypeptides lacking the barcode fragments of the first set of polypeptides; both the first set of polypeptides and the second set of polypeptides each comprise a reference fragment common to the first set of polypeptides and the second set of polypeptides and releasable by digestion with the protease; A mixture wherein the ratio of said first set of polypeptides to said polypeptide comprising said reference fragment is greater than 0.
70.
68. 68. The mixture of claim 67, wherein the ratio of the first set of polypeptides to the polypeptide comprising the reference fragment is greater than 0.8, 0.9, or 0.
95.
69. 69. The mixture of claim 67 or claim 68, wherein the reference fragment occurs twice in each polypeptide of the first set of polypeptides and the second set of polypeptides.
70. 68. The mixture of claim 67, wherein the first set of polypeptides comprises full-length polypeptides and the barcode fragments are portions of the full-length polypeptides.
71. 68. The mixture of claim 67, wherein the full-length polypeptide is a polypeptide of any one of claims 1 to 68.
72. 72. The mixture of any one of claims 70-71, wherein the barcode fragment does not include the N-terminal amino acid and the C-terminal amino acid of the full-length polypeptide.
73. 73. The mixture of any one of claims 67 to 72, wherein the mixture of polypeptides of varying lengths differ from each other due to N-terminal truncation, C-terminal truncation, or both N- and C-terminal truncation of full-length polypeptides.
74. A nucleic acid comprising a polynucleotide encoding a polypeptide according to any one of claims 1 to 75 or the reverse complement of said polynucleotide.
75. 75. An expression vector comprising the polynucleotide sequence of claim 74 and a control sequence operably linked to said polynucleotide sequence.
76. A host cell comprising the expression vector of claim 75.
77. 77. The host cell of claim 76, wherein the host cell is a prokaryote.
78. 78. The host cell of claim 77, wherein the host cell is Escherichia coli.
79. 79. The host cell of claim 78, wherein the host cell is a mammalian cell.
80. A pharmaceutical composition comprising a polypeptide according to any one of claims 1 to 73 and one or more pharmaceutically acceptable excipients.
81. Use of a polypeptide according to any one of claims 1 to 68 or a mixture thereof according to claims 67 to 73 in the preparation of a medicament for the treatment of a disease in a human or animal.
82. 82. The use of claim 81, wherein the disease is cancer.
83. 81. A method of treating a disease in a human or animal, comprising administering to said human or animal in need thereof one or more therapeutically effective doses of the pharmaceutical composition of claim 80.
84. 84. The method of claim 83, wherein the disease is cancer.
85. 84. The method of claim 83, wherein the human or animal is a human.
86. 1. A method for assessing the relative abundance of a first set of polypeptides in a mixture comprising polypeptides of varying lengths relative to a second set of polypeptides in the mixture, wherein each polypeptide in the first set of polypeptides shares a barcode fragment that occurs once in the polypeptides, each polypeptide in the second set of polypeptides lacks the barcode fragment shared by the polypeptides of the first set, and each individual polypeptide in both the first set of polypeptides and the second set of polypeptides comprises a reference fragment; contacting the mixture with a protease to generate a plurality of proteolytic fragments resulting from cleavage of the first set of polypeptides and the second set of polypeptides, wherein the plurality of proteolytic fragments comprises: a plurality of reference fragments; and contacting, the contact comprising a plurality of barcode fragments; and determining a ratio of the amount of the barcode fragment to the amount of the reference fragment, thereby assessing the relative amounts of the first set of polypeptides to the second set of polypeptides.
87. 87. The method of claim 86, wherein the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides.
88. 88. The method of claim 86 or 87, wherein the protease cleaves the polypeptides of variable length C-terminal to glutamic acid residues that are not followed by proline residues.
89. 89. The method of any one of claims 86 to 88, wherein the protease is Glu-C protease.
90. 90. The method of any one of claims 86 to 89, wherein the protease is not trypsin.
91. 91. The method of any one of claims 86-90, wherein determining the ratio of the amount of barcode fragments to the amount of reference fragments comprises quantifying barcode fragments and reference fragments from the mixture after contacting with the protease.
92. 92. The method of claim 91, wherein the barcode fragment and the reference fragment are identified based on their respective masses.
93. 93. The method of claim 91 or claim 92, wherein the barcode fragment and the reference fragment are identified via mass spectrometry.
94. 94. The method of any one of claims 91 to 93, wherein the barcode fragments and reference fragments are identified via liquid chromatography-mass spectrometry (LC-MS).
95. 95. The method of any one of claims 86 to 94, wherein determining the ratio of the barcode fragment to the reference fragment comprises isobaric labeling.
96. 96. The method of any one of claims 86-95, wherein determining the ratio of the barcode fragment to the reference fragment comprises adding to the mixture one or both of an isotopically labeled reference fragment and an isotopically labeled barcode fragment.
97. 97. The method of claim 96, wherein the barcode fragment, when present, is a portion of an XTEN polypeptide.
98. The method of any one of claims 86 to 97, wherein said mixture of polypeptides of different lengths comprises a polypeptide according to any one of claims 1 to 68.
99. 99. The method of any one of claims 86 to 98, wherein the polypeptides of various lengths include full-length polypeptides and truncated fragments thereof.
100. 100. The method of claim 99, wherein the polypeptides of various lengths are the full-length polypeptide and truncated fragments thereof.
101. 101. The method of any one of claims 86 to 100, wherein said mixture of polypeptides of different lengths differ from each other due to N-terminal truncations, C-terminal truncations, or both N- and C-terminal truncations of full-length polypeptides.
102. The method of claim 101, wherein the full-length polypeptide is a polypeptide according to any one of claims 1 to 68.
103. 103. The method of any one of claims 86 to 102, wherein the ratio of the amount of barcode fragment to reference fragment is greater than 0.5, 0.6, 0.7, 0.8, 0.9, or 0.
95.
104. A mixture comprising a plurality of polypeptides of various lengths, a first set of polypeptides, each polypeptide of the first set of polypeptides comprising a barcode fragment releasable from the polypeptide by digestion with a protease and having a sequence and molecular weight that differs from the sequences and molecular weights of all other fragments releasable from the first set of polypeptides; and a second set of polypeptides lacking the barcode fragments of the first set of polypeptides; each of the first set of polypeptides and the second set of polypeptides comprising a reference fragment common to the first set of polypeptides and the second set of polypeptides, the reference fragment being releasable by digestion with the protease; A mixture, wherein the number of reference fragments quantified in the polypeptide mixture after protease digestion is equal to the sum of the numbers of the first and second polypeptide sets in the mixture, and the number of barcode fragments quantified in the polypeptide mixture after protease digestion is equal to the number of the first polypeptide set in the mixture.
105. 105. The mixture of claim 104, wherein the first set of polypeptides comprises one reference fragment and the ratio of the first set of polypeptides to the polypeptides in the mixture comprising the reference fragment is greater than 0.
7.
106. 106. The mixture of claim 105, wherein the ratio of the first set of polypeptides to the polypeptide comprising the reference fragment is greater than 0.8, 0.9, or 0.
95.
107. 107. The mixture of any one of claims 104 to 106, wherein the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides.
108. 107. The mixture of any one of claims 104 to 106, wherein the reference fragment occurs twice in each polypeptide of the first set of polypeptides and the second set of polypeptides.
109. 107. The mixture of any one of claims 104 to 106, wherein the first set of polypeptides comprises full-length polypeptides and the barcode fragments are portions of the full-length polypeptides.
110. The mixture of claim 104 or claim 105, wherein the full-length polypeptide is a polypeptide according to any one of claims 1 to 66.
111. The barcode fragment is a fragment of the N-terminal amino acid and the C-terminal amino acid of the full-length polypeptide.
110. The mixture of claim 108 or claim 109, which is free of benzoic acid.
112. 112. The mixture of any one of claims 104 to 111, wherein said mixture of polypeptides of different lengths differ from each other due to N-terminal truncation, C-terminal truncation, or both N- and C-terminal truncation of full-length polypeptides.
113. 107. The mixture of any one of claims 104 to 106, wherein the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides.
114. 106. The mixture of claim 104 or claim 105, wherein the number of reference fragments in the first set of polypeptides may be different from the number of reference fragments in the second set of polypeptides, but the number in each polypeptide of each set must be the same.
115. 109. The mixture of claim 108, wherein each of the reference fragments in the polypeptides in the mixture has a sequence and molecular weight that differs from the sequence and molecular weight of all other fragments.
116. A mixture comprising a plurality of polypeptides of various lengths, A first set of polypeptides, wherein each polypeptide of said first set of polypeptides comprises: a first set of polypeptides comprising barcode fragments releasable from said polypeptides by digestion with a protease and having sequences and molecular weights that differ from the sequences and molecular weights of all other fragments releasable from said first set of polypeptides; and a second set of polypeptides lacking the barcode fragments of the first set of polypeptides; both the first set of polypeptides and the second set of polypeptides each comprise a reference fragment common to the first set of polypeptides and the second set of polypeptides and releasable by digestion with the protease; The ratio of the first set of polypeptides to the polypeptides in the mixture is determined by the equation [barcode-containing polypeptide] / [(reference peptide-containing polypeptide) x N], wherein N is the number of occurrences of said reference peptide released from each polypeptide in said mixture.
117. 117. The mixture of claim 116, wherein the first set of polypeptides comprises one reference fragment and the ratio of the first set of polypeptides to the polypeptides in the mixture comprising the reference fragment is greater than 0.
7.
118. 118. The mixture of claim 117, wherein the ratio of the first set of polypeptides to the polypeptide comprising the reference fragment is greater than 0.8, 0.9, or 0.
95.
119. 119. The mixture of any one of claims 116 to 118, wherein the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides.
120. 119. The mixture of any one of claims 116 to 118, wherein the reference fragment occurs twice in each polypeptide of the first set of polypeptides and the second set of polypeptides.
121. 117. The mixture of claim 115 or claim 116, wherein the first set of polypeptides comprises full-length polypeptides and the barcode fragments are portions of the full-length polypeptides.
122. The mixture of any one of claims 116 to 118, wherein the full-length polypeptide is a polypeptide of any one of claims 1 to 66.
123. 122. The mixture of claim 120 or claim 121, wherein the barcode fragment does not include the N-terminal amino acid and the C-terminal amino acid of the full-length polypeptide.
124. 124. The mixture of any one of claims 115 to 123, wherein said mixture of polypeptides of different lengths differ from each other due to N-terminal truncation, C-terminal truncation, or both N- and C-terminal truncation of full-length polypeptides.
125. 118. The mixture of any one of claims 115 to 117, wherein the reference fragment occurs no more than once in each polypeptide of the first set of polypeptides and the second set of polypeptides.
126. 117. The mixture of claim 115 or claim 116, wherein the number of reference fragments in the first set of polypeptides may be different from the number of reference fragments in the second set of polypeptides, but the number in each polypeptide of each set must be the same.
127. 121. The mixture of claim 120, wherein each of the reference fragments in the polypeptides in the mixture has a sequence and molecular weight that is different from the sequence and molecular weight of all other fragments.
128. 128. A method for detecting sequence integrity of polypeptides comprising a first set of polypeptides in a mixture of polypeptides according to claims 104-127, comprising the steps of digesting the mixture of polypeptides with a protease to release the barcode fragments and the reference fragments from the first set of polypeptides and to release the reference fragments from the second set of polypeptides, and determining a ratio of the barcode fragments from the first set of polypeptides to the reference fragments from the first and second sets of polypeptides, wherein the sequence integrity of polypeptides of the first set of polypeptides is detected by comparing the ratio of fragments to an expected ratio of the fragments based on the number of barcode fragments and reference fragments in polypeptides comprising the first and second sets of polypeptides.
129. 129. The method of claim 128, wherein the barcode fragment and the reference fragment are detected by LC / MS.
130. 130. The method of claim 129, wherein a known amount of a standard is added to the mixture of polypeptides, the standard comprising a plurality of isotopically labeled versions of the mixture of polypeptides of various lengths.
131. 131. The method of claim 130, wherein the isotopically labeled versions of the mixture of polypeptides of varying lengths are added to the mixture prior to digestion with the protease.
132. The method of claim 130, wherein the isotopically labeled versions of the mixture of polypeptides of various lengths are digested with the protease after the mixture has been digested with the protease and before being added to the mixture of claim 129.
133. 133. The method of any one of claims 130-132, further comprising quantitating the detected isotopically distinguishable amounts of the barcode fragment, the reference fragment, or both.
Citation Information
Patent Citations
Extended recombinant polypeptides and compositions comprising extended recombinant polypeptides
JP2012516854A
Methods for determining gene functions
US20180030531A1