Barcoded XTEN peptides and their compositions, as well as their preparation and use methods

By using XTEN peptide digestion technology containing non-overlapping sequence motifs and protease digestion, combined with LC-MS analysis, the problem of detecting and quantifying truncated variants in peptide mixtures has been solved, improving the biological properties and safety of the drug.

CN115175920BActive Publication Date: 2025-10-31AMUNIX PHARMACEUTICALS INC
View PDF 63 Cites 0 Cited by

Patent Information

Application Number
CN202080090841.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-13
Filing Date
2020-11-13
Publication Date
2025-10-31
Estimated Expiration
2040-11-13

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively identifying and quantifying truncated variants in peptide mixtures, leading to inconsistent biological behavior, affecting efficacy and safety, and existing methods have limited sensitivity.

Method used

XTEN peptides containing multiple non-overlapping sequence motifs were used to release specific barcode fragments via protease digestion for quantification and identification of truncated variants in peptide mixtures, followed by analysis using LC-MS technology.

Benefits of technology

This technology enables highly sensitive detection and quantification of truncated variants in peptide mixtures, improving the consistency and safety of drug biological properties and reducing the risk of immunogenicity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115175920B_ABST
    Figure CN115175920B_ABST
Patent Text Reader

Abstract

This document discloses polypeptides comprising an extended recombinant polypeptide (XTEN) containing multiple overlapping sequence motifs and one or more barcode fragments, which are releasable upon protease digestion and detectable from all other protease-releasable fragments. Some embodiments of these polypeptides also comprise a bioactive polypeptide, wherein advantageous embodiments include a protease-cleavable releasable segment capable of cleaving the link between the XTEN polypeptide and the bioactive polypeptide. Methods for preparing and using the polypeptides are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] sequence list

[0002] This application contains a sequence list, which has been electronically submitted in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy created on November 6, 2020, is named 20-1761-WO_Sequence_Listing_ST25.txt and is 1494 bytes in size. Background Technology

[0003] Peptides can be produced in a manner that results in a mixture of peptides. Peptide mixtures often include full-length peptides along with size variants (e.g., truncated ones). The presence of variants that differ in size from the desired full-length product can affect the biological behavior of the peptide drug substance, potentially impacting its safety and / or efficacy. For example, protein-based prodrugs for cancer treatment can be engineered to have a tumor-targeting activation mechanism. More specifically, full-length therapeutic proteins can be produced and administered in an inactive (non-cytotoxic) prodrug form, which is converted into an active drug by preferentially removing a portion of the prodrug peptide at the intended biological site (e.g., tumor). Truncated variants of the full-length construct can lose their protective sequence and become cytotoxic (active), thereby “contaminating” the prodrug composition and producing a mixture with components that are not intentionally active outside the intended biological site. In some cases, such shorter variants can pose a greater risk of immunogenicity, exhibit less selective toxicity to tumor cells, or display a less desirable pharmacokinetic profile compared to the full-length protein (e.g., resulting in a narrower therapeutic window), or harmfully have unintended effects on the receptor outside the intended site (e.g., in healthy tissue). Consequently, the detection and quantification of protein structural changes can be important for evaluating the biological properties of biotherapeutic agents (e.g., clinical safety and pharmacological efficacy) and for developing new biotherapeutic agents (e.g., with increased efficacy and reduced side effects). Existing techniques and methods for identifying and quantifying the amount of “contaminated” truncated products may include one or more drawbacks, such as limited sensitivity, ease of use, efficiency, or effectiveness. Summary of the Invention

[0004] This document discloses polypeptides comprising extended recombinant polypeptides (XTENs), said extended recombinant polypeptides comprising a plurality of non-overlapping sequence motifs. In the XTEN polypeptide of the present invention, the plurality of non-overlapping sequence motifs comprises: a set of non-overlapping sequence motifs, each of which is repeated at least twice in the XTEN polypeptide; and a unique non-overlapping sequence motif that appears only once within the XTEN polypeptide; said polypeptide further comprises a first barcode fragment that can be released from the polypeptide upon digestion by a protease. In the embodiments described herein, the first barcode fragment is part of the XTEN, comprising at least a portion of the sequence motif that appears only once within the XTEN, and is different in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide upon complete digestion by a protease. Further, in the XTEN embodiments of the present invention provided herein, the barcode fragment does not include an N-terminal amino acid or a C-terminal amino acid of the polypeptide. As further disclosed herein, the XTEN polypeptide of the present invention is characterized by comprising a length of at least 150 amino acids, more specifically, a length of 150-3000 amino acids. The amino acids constituting the XTEN polypeptide of the present invention are characterized, wherein at least 90% of these residues are glycine (G), alanine (A), serine (S), threonine (T), glutamate (E), or proline (P), and the XTEN polypeptide contains at least four of these amino acids (G, A, S, T, E, or P). Additionally, the XTEN polypeptide provided herein contains non-overlapping sequence motifs of 9 to 14 amino acids in length, and within each of these non-overlapping motifs, the sequences having G, A, S, T, E, or P amino acids are substantially randomized with respect to any other non-overlapping sequence motifs constituting the XTEN polypeptide.

[0005] In some embodiments, the barcode fragment does not include a glutamic acid immediately adjacent to another glutamic acid in XTEN. In some embodiments, the barcode fragment has a glutamic acid at its C-terminus. In some embodiments, the barcode fragment has an N-terminal amino acid immediately preceding a glutamic acid residue. In some embodiments, the glutamic acid residue preceding the N-terminal amino acid is not immediately adjacent to another glutamic acid residue. In some embodiments, the barcode fragment does not include glutamic acid residues at positions other than the C-terminus of the barcode fragment, unless the glutamic acid is immediately followed by proline. In some embodiments, the barcode fragment is positioned 10 to 150 amino acids away from the N-terminus or C-terminus of the polypeptide.

[0006] In some embodiments, the sequence motifs of this set of non-overlapping sequence motifs are identified herein by SEQ ID NO: 182-203 and 1715-1722. In some embodiments, the sequence motifs of this set of non-overlapping sequence motifs are identified herein by SEQ ID NO: 186-189. In some embodiments, the set of non-overlapping sequence motifs comprises at least two, at least three, or all four of the sequence motifs SEQ ID NO: 186-189.

[0007] In specific embodiments, the polypeptides provided herein comprise XTEN polypeptides as disclosed herein, wherein the barcode fragments do not include the N-terminal or C-terminal amino acids of the polypeptide; do not include glutamic acid immediately adjacent to another glutamic acid in XTEN; have glutamic acid at its C-terminus; have an N-terminal amino acid immediately preceding a glutamic acid residue; and are located 10 to 125 amino acids from the N-terminus or C-terminus of the polypeptide.

[0008] In some of these specific embodiments, the glutamate residue preceding the N-terminal amino acid is not immediately adjacent to another glutamate residue. In some of these specific embodiments, the barcode fragment does not include glutamate residues at positions other than the C-terminus of the barcode fragment, unless the glutamate is immediately followed by proline.

[0009] In some embodiments, the XTEN peptide provided herein comprises a plurality of non-overlapping sequence motifs, wherein each of the sequence motifs is repeated at least twice in the XTEN peptide and has a length of 9 to 14 amino acids. In some embodiments, the sequence motifs of this set of non-overlapping sequence motifs are identified herein by SEQ ID NO: 182-203 and 1715-1722. In some embodiments, the sequence motifs of this set of non-overlapping sequence motifs are identified herein by SEQ ID NO: 186-189. In some embodiments, the set of non-overlapping sequence motifs comprises at least two, at least three, or all four of the sequence motifs SEQ ID NO: 186-189. In some embodiments, at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid residues of the XTEN polypeptide are combinations of glycine (G), alanine (A), serine (S), threonine (T), glutamate (E), or proline (P), wherein the XTEN polypeptide contains at least four of these amino acids (G, A, S, T, E, or P). In some embodiments, the XTEN is 150 to 3000 amino acids in length. In some embodiments, the XTEN is 150 to 1000 amino acids in length. In some embodiments, the polypeptide can be cleaved by a protease that cleaves at the C-terminal side of a glutamate residue, which is subsequently not proline. In some embodiments, the protease is a Glu-C protease.

[0010] In some embodiments of the XTEN peptides provided herein, the barcode fragment is located within 200, 150, 100, or 50 amino acids from the N-terminus of the peptide. In some embodiments, the barcode fragment is located between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the N-terminus of the protein. In some embodiments, the barcode fragment is located within 200, 150, 100, or 50 amino acids from the C-terminus of the peptide. In some embodiments, the barcode fragment is located between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the C-terminus of the protein. In some embodiments, the length of the barcode fragment is at least 4 amino acids. In some embodiments, the length of the barcode fragment is 4 to 20, 5 to 15, 6 to 12, or 7 to 10 amino acids. In some embodiments, the barcode fragment is identified herein by SEQ ID No: 8020-8030 (BAR001-BAR011).

[0011] In some embodiments, the polypeptide further comprises a second barcode fragment, wherein the second barcode fragment is part of XTEN and is different in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide after complete digestion by the protease. In some embodiments, the polypeptide further comprises a third barcode fragment, wherein the third barcode fragment is part of XTEN and is different in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide after complete digestion by the protease.

[0012] In some embodiments, XTEN has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with the sequence identified herein by SEQ ID NO:8001-8019. In some embodiments, the length of XTEN is at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids.

[0013] In some embodiments, the polypeptide further comprises a bioactive polypeptide linked to an XTEN polypeptide (BPXTEN). In some embodiments, the XTEN polypeptide is linked to the bioactive polypeptide at an amino or carboxyl terminus of the XTEN. In any configuration, the barcode fragment is located within a region of the XTEN that extends 5% to 50%, 7% to 40%, or 10% to 30% of the length of the XTEN, as measured from the amino or carboxyl terminus linked to the bioactive polypeptide.

[0014] In some embodiments, the BPXTEN polypeptide further includes one or more reference fragments that are releasable from the polypeptide upon digestion by a protease, each of the reference fragments comprising a portion of the bioactive polypeptide. In some embodiments, the one or more reference fragments are single reference fragments that are different in sequence and molecular weight from all other peptide fragments that are releasable from the polypeptide upon digestion by a protease. In some embodiments, the reference fragment comprises a peptide whose presence in the polypeptide mixture indicates its presence or integrity (i.e., the protein has not yet been degraded or cleaved by a protease).

[0015] In some embodiments, the BPXTEN polypeptide further comprises a first release region (RS1) located between the XTEN and the bioactive polypeptide. In some embodiments, RS1 comprises an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a sequence identified herein by any of the sequences in Tables 4a-4h and 8a-8b. In some embodiments, the bioactive polypeptide is identified herein by any of the sequences or combinations thereof in Tables 4a-4h and 8a-8b.

[0016] In some embodiments, BPXTEN peptides advantageously have at least twice the terminal half-life compared to bioactive peptides not linked to any XTEN.

[0017] In some embodiments, BPXTEN peptides are advantageously less immunogenic than bioactive peptides not linked to any XTEN, wherein immunogenicity can be determined by measuring the production of IgG antibodies selectively bound to the bioactive peptide after administration of a comparable dose to a human or animal.

[0018] In some embodiments, the BPXTEN peptide exhibits an apparent molecular weight factor greater than about 6 under physiological conditions.

[0019] In some embodiments, the BPXTEN polypeptide further comprises a second XTEN polypeptide, wherein the second XTEN polypeptide comprises an amino acid sequence having the same characteristics as the first XTEN component described throughout the foregoing and present disclosure of these embodiments of BPXTEN, and wherein the first XTEN polypeptide is located at the N-terminus of the bioactive polypeptide, and the second XTEN polypeptide is located at the C-terminus of the bioactive polypeptide. In some embodiments, the second XTEN polypeptide comprises an amino acid sequence different from the amino acid sequence of the first XTEN constituting these embodiments of BPXTEN. In some embodiments, the amino acid sequence of the second XTEN polypeptide is longer than the amino acid sequence of the first XTEN polypeptide.

[0020] In some embodiments, the BPXTEN peptide further comprises a second release segment (RS2) located between the bioactive peptide and the second XTEN peptide. In some embodiments, the RS1 and RS2 sequences of the first XTEN peptide are identical. In some embodiments, the RS1 of the first XTEN peptide and the RS2 of the second XTEN peptide are each substrates for cleavage by multiple proteases at one, two, three, or more cleavage sites within each release segment sequence.

[0021] In some of these embodiments, the BPXTEN polypeptide includes a further barcode fragment that is part of the second XTEN polypeptide and is different in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide after complete digestion by the protease. In some of these embodiments, the further barcode fragment does not include the C-terminal amino acid of the polypeptide. In some of these embodiments, the further barcode fragment includes a glutamic acid residue at its C-terminus. In some of these embodiments, the further barcode fragment of the second XTEN polypeptide is located within 200, 150, 100, or 50 amino acids from the C-terminus of the second XTEN component of the BPXTEN polypeptide. In some of these embodiments, the further barcode fragment of the second XTEN polypeptide is located at a position between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the C-terminus of the second XTEN component of the BPXTEN polypeptide. In some of these embodiments, the length of the further barcode fragment is 4 to 20, 5 to 15, 6 to 12, or 7 to 10 amino acids. In some of these embodiments, further barcode fragments are identified herein by SEQ ID No: 8020-8030 (BAR001-BAR011).

[0022] In some embodiments, the second XTEN polypeptide further comprises a set of barcode fragments, including further barcode fragments and at least one additional barcode fragment, wherein each barcode fragment in the set is different in sequence and molecular weight from all other peptide fragments that can be released from the BPXTEN polypeptide after complete digestion of the polypeptide by the protease. In some embodiments, the second XTEN polypeptide is identified by SEQ ID NO:8001-8019. In some embodiments, the further barcode fragments do not include a glutamic acid residue immediately adjacent to another glutamic acid residue in the polypeptide.

[0023] In some embodiments, at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid residues of the second XTEN polypeptide are combinations of glycine (G), alanine (A), serine (S), threonine (T), glutamate (E), and proline (P), wherein the XTEN polypeptide contains at least four of these amino acids (G, A, S, T, E, or P). In some embodiments, the total number of amino acids in the first XTEN polypeptide and the sum of the total number of amino acids in the second XTEN polypeptide is at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, or at least 800 amino acids. In some embodiments, the second XTEN polypeptide contains a plurality of non-overlapping sequence motifs, each of which is repeated at least twice in the second XTEN polypeptide sequence and has a length of 9 to 14 amino acids.

[0024] In some embodiments, for the second XTEN polypeptide, the sequence motifs of a plurality of non-overlapping sequence motifs are identified herein by SEQ ID NO: 182-203 and 1715-1722. In some embodiments, the sequence motifs of a plurality of non-overlapping sequence motifs are identified herein by SEQ ID NO: 186-189. In some embodiments, for the second XTEN polypeptide, the plurality of non-overlapping sequence motifs comprise at least two, at least three, or all four of the following motifs: SEQ ID NO: 186-189. In some embodiments, the second XTEN polypeptide is 150 to 3000 amino acids in length. In some embodiments, the second XTEN polypeptide is 150 to 1000 amino acids in length. In some embodiments, the second XTEN polypeptide has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with the sequences identified herein by SEQ ID NO: 8001-8019. In some embodiments, the length of the second XTEN polypeptide is at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids.

[0025] In certain embodiments, the BPXTEN peptides provided herein comprise a first XTEN peptide containing a first RS sequence adjacent to, but not constituting, the C-terminus of the peptide, the first XTEN peptide being covalently linked to a tandemly covalently linked first and second bioactive peptides, wherein the second XTEN peptide is covalently linked to the C-terminus of the tandemly linked bioactive peptides, wherein the second XTEN peptide contains a second RS sequence adjacent to, but not constituting, the N-terminus of the second XTEN peptide, wherein the first RS sequence and the second RS sequence may be the same or different. In certain embodiments, the second XTEN peptide comprises an amino acid sequence longer than the amino acid sequence of the first XTEN peptide. In some embodiments, the first bioactive protein or the second bioactive protein, or both, comprise a specific binding protein, wherein in some embodiments, the specific binding protein specifically binds to an antigen or agonist expressed at a desired biological site. In certain embodiments, the desired biological site is a tumor, and the antigen is a tumor-specific antigen. In certain embodiments, the first bioactive peptide and the second bioactive peptide are different, including but not limited to having different specific binding affinities.

[0026] This document also discloses nucleic acids comprising polynucleotides encoding polypeptides such as any XTEN or BPXTEN polypeptide disclosed herein, or the reverse complement of said polynucleotides.

[0027] This document also discloses expression vectors comprising any polynucleotide sequence disclosed herein and a regulatory sequence operatively linked to the polynucleotide sequence, the regulatory sequence regulating the expression or other biological activities of the polynucleotide sequence.

[0028] This document discloses host cells comprising expression vectors as disclosed herein. In some embodiments, the host cell is a prokaryote. In some of these embodiments, the host cell is *Escherichia coli*. In some alternative embodiments, the host cell is a mammalian cell.

[0029] This document also discloses pharmaceutical compositions comprising polypeptides as disclosed herein and one or more pharmaceutically acceptable excipients. In some embodiments, the pharmaceutical compositions are formulated for administration to animals, and particularly humans, wherein said administration may be by any therapeutically effective route of administration. Pharmaceutical compositions disclosed herein can be prepared and used in any formulation known in the art and particularly suitable for the route of administration, site of administration, and intended effect on humans or animals.

[0030] This document discloses the use of polypeptides as disclosed herein, and in particular BPXTEN polypeptides, in the preparation of medicaments for treating diseases, conditions, or illnesses in humans or animals. In some embodiments, the disease, condition, or illness may be cancer.

[0031] This document discloses a method for treating diseases in humans or animals, as described above and throughout this disclosure, comprising administering one or more therapeutically effective doses of a pharmaceutical composition to a human or animal in need. In some embodiments, the pharmaceutical composition is administered to a human or animal as one or more therapeutically effective doses, said therapeutically effective doses being administered once daily, weekly, monthly, or annually on a clinically appropriate schedule and at a clinically appropriate dosage.

[0032] This document discloses mixtures comprising various peptides of varying lengths, particularly XTEN and BPXTEN peptides as disclosed herein, wherein the mixture comprises:

[0033] A first group of polypeptides, wherein each polypeptide in the first group of polypeptides comprises a barcode fragment, the barcode fragment being releaseable from the polypeptide by digestion with a protease, and having a sequence and molecular weight different from the sequence and molecular weight of all other fragments that can be released from the first group of polypeptides; and

[0034] A second group of peptides lacking the barcode fragments of the first group of peptides;

[0035] The first group of polypeptides and the second group of polypeptides each contain a reference fragment, which is common to both groups and is produced by digestion with a protease; and

[0036] The ratio of the first group of polypeptides to the polypeptides containing the reference fragment is greater than 0.7.

[0037] In some embodiments, the ratio of the first group of peptides to the peptide containing the reference fragment is greater than 0.8, 0.9, 0.95, or 0.98. In some embodiments, the reference fragment appears no more than once in each of the first and second group of peptides. In some embodiments, the protease is a protease that cleaves at the C-terminus of a glutamate residue. In some embodiments, barcode release from the peptide containing the first group of peptides is promoted by pepsin, elastase, thermophilic protease, or Glu-C protease. In some embodiments, barcode release is promoted by Glu-C protease. In some embodiments, the protease is not trypsin. In some embodiments, peptides of various lengths include peptides containing at least one XTEN peptide as described herein.

[0038] In some embodiments, the first group of peptides comprises a full-length peptide, wherein the barcode fragment is a portion of the full-length peptide. In some embodiments, the full-length peptide is any peptide disclosed herein, and particularly XTEN and BPXTEN peptides. In some embodiments, the barcode fragment does not contain an N-terminal or C-terminal amino acid of the full-length peptide. In some embodiments, mixtures of peptides of various lengths differ from each other due to N-terminal truncation, C-terminal truncation, or both N-terminal and C-terminal truncation of the full-length peptide.

[0039] This document discloses a method for evaluating the relative amounts of a first group of peptides and a second group of peptides in a mixture comprising peptides of various lengths, particularly XTEN and BPXTEN peptides as disclosed herein, wherein each peptide in the first group shares a barcode fragment that appears once and only once in the peptide, and each peptide in the second group lacks a barcode fragment shared by the first group of peptides, wherein each peptide in both the first and second groups of peptides comprises a reference fragment, the method comprising:

[0040] The mixture is contacted with a protease to produce multiple protease-digested fragments derived from the cleavage of a first group of polypeptides and a second group of polypeptides, wherein the multiple protease-digested fragments comprise multiple reference fragments and multiple barcode fragments; and

[0041] The ratio of the amount of the barcode fragment to the amount of the reference fragment is determined to evaluate the relative amounts of the first group of peptides and the second group of peptides.

[0042] In some embodiments, the reference fragment appears no more than once in each of the first group of peptides and the second group of peptides.

[0043] In some embodiments, the protease cleaves polypeptides of various lengths at the C-terminal side of a glutamate residue, which is subsequently not a proline residue. In some embodiments, the protease is a Glu-C protease. In some embodiments, the protease is not a trypsin. In some embodiments, determining the ratio of the amount of barcode fragment to the amount of reference fragment includes quantifying the barcode fragment and reference fragment from the mixture after the polypeptide mixture has been contacted with the protease. In some embodiments, the barcode fragment and reference fragment are identified based on their respective masses. In some embodiments, the barcode fragment and reference fragment are identified by mass spectrometry. In some embodiments, the barcode fragment and reference fragment are identified by liquid chromatography-mass spectrometry (LC-MS). In some embodiments, determining the barcode fragment / reference fragment ratio includes isotopic labeling or stable isotope labeling. In some embodiments, determining the barcode fragment / reference fragment ratio includes one or a mixture of isotopically labeled reference fragments and isotopically labeled barcode fragments.

[0044] In some of these embodiments, polypeptides of various lengths comprise full-length polypeptides and truncated fragments thereof. In some of these embodiments, mixtures of polypeptides of various lengths differ from one another due to N-terminal truncation, C-terminal truncation, or both N-terminal and C-terminal truncation of the full-length polypeptide. In some of these embodiments, the ratio of barcode fragment to reference fragment is greater than 0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.98, or 0.99.

[0045] This document discloses a mixture comprising multiple peptides of various lengths, the mixture comprising a first group of peptides, wherein each peptide in the first group comprises a barcode fragment that can be released from the peptide by digestion with a protease and has a sequence and molecular weight different from that of all other fragments that can be released from the first group of peptides. The embodiments also include a second group of peptides lacking barcode fragments from the first group of peptides, wherein both the first and second groups of peptides each comprise a reference fragment common to both groups and can be released by digestion with a protease. In the embodiments, the number of reference fragments quantified in the peptide mixture after protease digestion is equal to the sum of the number of the first and second groups of peptides in the mixture, and the number of barcode fragments quantified in the peptide mixture after protease digestion is equal to the number of the first group of peptides in the mixture. In the embodiments, the first group of peptides comprises one reference fragment, and the ratio of the first group of peptides to peptides containing the reference fragment in the mixture is greater than 0.7.

[0046] In some embodiments, the mixture has a ratio of a first group of peptides to a peptide containing a reference fragment greater than 0.8, 0.9, or 0.95.

[0047] In one particular embodiment, the reference fragment appears no more than once in each of the first and second groups of peptides. In an alternative embodiment, the reference fragment appears twice in each of the first and second groups of peptides.

[0048] In some embodiments, the first group of polypeptides comprises a full-length polypeptide, wherein the barcode fragment is a portion of the full-length polypeptide.

[0049] In some embodiments, the full-length polypeptide includes the polypeptides disclosed herein.

[0050] In one particular embodiment, the mixture barcode fragment does not contain the N-terminal and C-terminal amino acids of the full-length polypeptide.

[0051] In some embodiments, the mixture contains polypeptides of various lengths, which differ from one another due to N-terminal truncation, C-terminal truncation, or both N-terminal and C-terminal truncation of the full-length polypeptide.

[0052] In some embodiments, the reference fragment appears no more than once in each of the first and second groups of peptides. In an alternative embodiment, the number of reference fragments in the first group of peptides may differ from the number of reference fragments in the second group of peptides, but the number of reference fragments in each peptide within each group must be the same.

[0053] In one particular embodiment, the reference fragments in the polypeptide mixture each have a sequence and molecular weight that differ from the sequences and molecular weights of all other fragments.

[0054] This document discloses a mixture comprising multiple peptides of various lengths, the mixture comprising a first group of peptides, wherein each peptide in the first group comprises

[0055] A barcode fragment, which can be released from the polypeptide by digestion with a protease, and has a sequence and molecular weight different from that of all other fragments that can be released from the first group of polypeptides. The mixture also contains a second group of polypeptides lacking the barcode fragment from the first group of polypeptides, wherein both the first and second groups of polypeptides each contain a reference fragment common to both groups and can be released by digestion with a protease. The ratio of the first group of polypeptides to the polypeptides in the mixture has the following formula:

[0056] [Polypeptide containing barcode] / [(Polypeptide containing reference peptide) x N]

[0057] Where N is the number of times a reference peptide is released from each polypeptide in the mixture, and where the ratio of the first group of polypeptides containing a reference fragment to the polypeptides containing the reference fragment is greater than 0.7 when the first group of polypeptides contains a reference fragment.

[0058] In one particular embodiment, the ratio of the first group of peptides to the peptide containing the reference fragment is greater than 0.8, 0.9, or 0.95.

[0059] In some embodiments, the reference fragment appears no more than once in each of the first group of peptides and the second group of peptides.

[0060] In some embodiments, the reference fragment appears twice in each of the first group of peptides and the second group of peptides.

[0061] In one particular embodiment, the first group of polypeptides comprises a full-length polypeptide, wherein the barcode fragment is a portion of the full-length polypeptide.

[0062] In some embodiments, the full-length polypeptide includes the polypeptides disclosed herein. In one particular embodiment, the barcode fragment does not contain the N-terminal and C-terminal amino acids of the full-length polypeptide.

[0063] In some embodiments, mixtures of peptides of various lengths are different from each other due to truncation of the N-terminus, C-terminus, or both of the N-terminus and C-terminus of the full-length peptide.

[0064] In some embodiments, the reference fragment appears no more than once in each of the first and second groups of peptides. In further embodiments, the number of reference fragments in the first group of peptides may differ from the number of reference fragments in the second group of peptides, but the number of reference fragments in each peptide within each group must be the same. In some embodiments, the reference fragment in the peptides of the mixture has a sequence and molecular weight that differ from the sequence and molecular weight of all other fragments.

[0065] This document discloses a method for detecting the sequence integrity of a polypeptide comprising a first group of polypeptides in a mixture disclosed herein. The method includes the step of digesting the polypeptide mixture with a protease, the protease releasing barcode fragments and reference fragments from the first group of polypeptides and releasing reference fragments from the second group of polypeptides, and determining the ratio of barcode fragments from the first group of polypeptides to reference fragments from the first and second groups of polypeptides. In one particular embodiment, the sequence integrity of the polypeptide comprising the first group of polypeptides is detected by comparing the ratio of fragments to an expected ratio of fragments based on the number of barcode fragments and reference fragments in the polypeptide comprising the first and second groups of polypeptides.

[0066] The methods considered herein readily accommodate, for example, the qualitative and quantitative analysis of peptides containing barcode and / or reference fragments using LC / MS. In one particular embodiment, LC / MS is quantitative and detects isotopically distinguishable amounts of the barcode fragment, reference fragment, or both. In exemplary such methods, the mixture of peptides is doped with a known amount of a “standard material” to facilitate such analysis. For example, such a standard material is a standard material in the form of an isotopically labeled form comprising a mixture of said various lengths of peptides to be analyzed. This isotopically labeled standard can be added to the mixture as a complete sequence prior to digestion by the protease. Alternatively, the test sample of the mixture of peptides of various lengths and the isotopically labeled standard material are digested by the protease in separate reactions, and the protease-digested isotopically labeled standard material is added to the test sample prior to analysis by LC / MS. The method of the present invention also includes quantifying the amount of the barcode fragment, reference fragment, or both from the test sample by comparison with the quantification of the detected isotopically distinguishable amount of the barcode fragment, reference fragment, or both.

[0067] Upon review of this disclosure, those skilled in the art will recognize variations and modifications to these embodiments. The foregoing features and aspects can be implemented in any combination and sub-combination (including multiple dependent combinations and sub-combinations) of one or more other features described herein. The various features described or illustrated above, including any components thereof, can be combined or integrated in other embodiments. Furthermore, certain features may be omitted or not implemented.

[0068] Incorporate by reference

[0069] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the same extent that each individual publication, patent or patent application is specifically and individually indicated to be incorporated by reference. Attached Figure Description

[0070] Various features of this disclosure are particularly set forth in the appended claims. A better understanding of the features and advantages of this disclosure can be obtained by referring to the following detailed description, along with the accompanying drawings, which illustrate illustrative embodiments utilizing the principles of the invention, in which:

[0071] Figure 1This describes a mixture of XTENylated protease-activated T cell engager (“XPAT”) peptides with various lengths. The full-length XPAT (top) comprises a 288-amino acid XTEN peptide at the N-terminus and an 864-amino acid XTEN peptide at the C-terminus. Various truncations can occur in one or both of the N-terminal and C-terminal XTEN peptides in the XPAT, for example, during fermentation, purification, or other steps in product preparation. While products with limited truncation (truncation close to the distal portion of the XTEN peptide from which it is linked) can function in a manner similar to the full-length construct, severely truncation (truncation closer to the proximal portion of the XTEN peptide from which it is linked) can have significantly different pharmacological properties than their full-length counterparts. The presence of truncation poses a challenge to the quantification of pharmacologically effective and ineffective variants in XPAT products. Figure 1 As shown in full-length XPAT, each XTEN peptide has a proximal and a distal end, wherein the proximal end is positioned closer to the bioactive peptide (e.g., T cell conjugates, cytokines, monoclonal antibodies (mAbs), antibody fragments, or other XTEN-modified proteins) than the distal end. Depending on the bonding orientation, the proximal or distal end of the XTEN peptide may correspond to the N-terminus or C-terminus of the XTEN peptide.

[0072] Figure 2A mixture of XPAT peptides with barcoded XTEN peptides of various lengths was depicted. In the full-length XPAT (top), the 288-amino acid-long N-terminal XTEN peptide contains three cleavably fused barcode sequences, “NA”, “NB”, and “NC” (from distal to proximal), and the 864-amino acid-long C-terminal XTEN peptide contains three cleavably fused barcode sequences, “CC”, “CB”, and “CA” (from proximal to distal). Each barcode is positioned to indicate the pharmacologically relevant length of the corresponding XTEN peptide. For example, a minor N-terminal truncated product of XPAT lacking barcode “NA” but having the more proximal barcodes “NB” and “NC” can exhibit substantially the same pharmacological properties as the full-length construct. In contrast, a major N-terminal truncated product of XPAT lacking, for example, all three barcodes at the N-terminus can be identifiable as different from the full-length construct in terms of pharmacological activity. A unique proteolytic cleavable sequence was identified from the bioactive polypeptide of XPAT (here, the tandem scFv containing the active portion of the T-cell binding agent). Because it is present in all length variants of XPAT (including full-length XPAT, minor truncated, and major truncated), the unique proteolytic cleavable sequence can be used as a reference for quantifying the amount of various truncated products relative to the total amount of bioactive protein.

[0073] Figure 3 This illustrates a potential design for a barcoded XTEN peptide by inserting a barcode-generating sequence into a generic (or conventional) XTEN peptide. An exemplary generic (or conventional) XTEN peptide (top) contains a non-overlapping 12-mer motif in the sequence “BCDABDCDABDCBDCDABDCB”, where the sequence motifs “A”, “B”, “C”, and “D” appear 3, 6, 5, and 7 times, respectively. Glu-C protease digestion of the exemplary generic XTEN peptide (top) does not produce a unique peptide other than the two ends (“NT” and “CT”). Inserting a barcode-generating sequence “X” (e.g., a unique 12-mer) into the XTEN peptide results in a unique proteolytically cleavable sequence (or barcode sequence) that does not appear anywhere else in the XTEN peptide. The barcode-generating sequence “X” can be positioned such that the resulting barcode marks the pharmacologically relevant length of the XTEN peptide. For example, an XTEN peptide lacking a barcode can be functionally different from a corresponding XTEN peptide with a barcode. Those skilled in the art will understand that the barcode generation sequence (“X”) can be the barcode sequence itself. Alternatively, the barcode generation sequence (“X”) can be different from the resulting barcode sequence. For example, the barcode sequence can overlap with and thus contain portions of a preceding or following 12-mer motif.

[0074] Figures 4A-4BThe quantification of the truncation level of the N-terminal XTEN peptide is shown. Figure 4A It has been confirmed that a barcoded XTEN peptide (bottom figure) can be constructed by replacing the sequence motif (e.g., the third sequence motif from the N-terminus, "D") in a generic XTEN peptide (top figure) with a barcode-generating motif "X"; and in this example, the barcode-generating motif ("X") is itself a uniquely proteolytically cleavable barcode sequence. Figure 4A As shown in the diagram below, the barcodes are positioned such that all severely truncated forms of the XTEN peptide lack barcodes, while all finitely truncated forms of the XTEN peptide contain barcodes. Figure 4B The relative abundance of various cleavage products in two different mixtures of XPAT is shown. In one mixture, barcodes were present in 99% of the constructs containing bioactive proteins. In the other mixture, 13% of the constructs lacked barcodes. Figures 4A-4B The use of barcoded XTEN peptides to distinguish two peptide mixtures with substantially similar average molecular weights but distinctly different pharmacological activities is illustrated.

[0075] Figure 5A This paper illustrates the analytical size exclusion chromatography (SEC) method for XPAT protein and the detection of the full-length protein and its truncated derivatives. The synthetic protein + truncated fraction includes fragments as large as the full-length synthetic protein.

[0076] Figure 5B The abundance of barcode peptides in XPAT formulations, as detected by mass spectrometry, is shown. Each measurement is the XIC area of ​​the N-barcode SGPGSTPAE (SEQ ID No. 8029) and the C-barcode GSAPGTE (SEQ ID No. 8023), normalized to the 400 nM peak of their respective heavy isotope-labeled synthetic peptide.

[0077] The patent or application documents contain at least one color drawing. A copy of the color drawing disclosed in this patent or patent application will be provided by the Patent Office upon request and payment of the necessary fees.

[0078] the term

[0079] As used herein, unless otherwise stated, the following terms have their respective meanings.

[0080] As used in the specification and claims, the singular forms “an,” “a,” and “the” include the plural referents unless the context clearly indicates otherwise. For example, the term “cell” includes a variety of cells, including mixtures thereof.

[0081] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to amino acid polymers of any length. Polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acid components. The term also covers amino acid polymers that have been modified, for example, by disulfide bond formation, glycosylation, esterification, acetylation, phosphorylation, or any other operation such as conjugation with a labeled component.

[0082] As used herein, the term "amino acid" refers to natural and / or non-natural or synthetic amino acids, including but not limited to glycine and both its D or L optical isomers, as well as amino acid analogs and peptides. Standard single-letter or three-letter codes are used to designate amino acids.

[0083] "Host cell" includes an individual cell or cell culture that can be, or has been, a recipient of a human or animal vector. Host cells include the offspring of a single host cell. Due to naturally occurring or genetically modified variations, the offspring are not necessarily completely identical to the original parent cell (in terms of morphology or the genome of total DNA complementarity).

[0084] A "chimeric" protein contains at least one polypeptide comprising a region in its sequence located at a different position than that found in nature. These regions may normally be present in separate proteins and are bound together in a fusion polypeptide; or they may normally be present in the same protein but are arranged in a new arrangement in the fusion polypeptide. The protein may be described as a "conjugated," "linked," "fused," or "fusion" protein; these terms are used interchangeably herein and refer to the linking of two or more polypeptide sequences together by any means, including chemical conjugation or recombination. For example, chimeric proteins can be produced by chemical synthesis or by generating and translating polynucleotides in which peptide regions are encoded in a desired relationship.

[0085] The terms "polynucleotide," "nucleic acid," "nucleotide," and "oligonucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, which are deoxyribonucleotides or ribonucleotides or analogues thereof. Polynucleotides can have any three-dimensional structure and can perform any known or yet-to-be-discovered or developed function. Polynucleotides may contain modified nucleotides, such as methylated nucleotides and nucleotide analogues. If present, modifications to the nucleotide structure can be conferred before or after polymer assembly. The nucleotide sequence can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, for example, by conjugation with labeled components.

[0086] The term "polynucleotide complement" refers to a polynucleotide molecule that has a complementary base sequence and reverse orientation compared to a reference sequence, in which it can hybridize with the reference sequence with perfect fidelity.

[0087] As used herein, a polynucleotide having “homology” or “homologous” is a polynucleotide that hybridizes to those sequences under strict conditions as defined herein and has at least 70%, preferably at least 80%, more preferably at least 90%, more preferably 95%, more preferably 97%, more preferably 98%, and even more preferably 99% sequence identity with those sequences.

[0088] When applied to polynucleotide sequences, the terms "percentage identity" and "% identity" refer to the percentage of residue matches between at least two polynucleotide sequences aligned using a normalization algorithm. Such algorithms can insert gaps in the sequences being compared in a normalized and reproducible manner to optimize the alignment between the two sequences and thus enable a more meaningful comparison. Percentage identity can be measured, for example, over the entire length of the defined polynucleotide sequence as defined by a specific SEQ ID number, or over a shorter length, such as over the length of a fragment taken from a larger defined polynucleotide sequence, for example, a fragment of at least 45, at least 60, at least 90, at least 120, at least 150, at least 210, or at least 450 consecutive residues. Such lengths are merely exemplary, and it should be understood that any fragment length supported by the sequences shown herein in tables, figures, or sequence listings can be used to describe lengths on which percentage identity can be measured.

[0089] The "percentage (%) amino acid sequence identity" of the polypeptide sequence identified herein is defined as the percentage of amino acid residues in the query sequence that are identical to amino acid residues in a second reference polypeptide sequence or a portion thereof, after alignment and, if necessary, the introduction of vacancies to achieve maximum percentage sequence identity, without considering any conserved substitutions as part of the sequence identity. Alignment for determining percentage amino acid sequence identity can be performed in various ways within the art, such as using publicly available computer software like BLAST, BLAST-2, ALIGN, or Megalign (DNASTAR). Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithm required to achieve maximum alignment across the full length of the sequences being compared. Percentage identity can be measured, for example, over the length of the entire defined polypeptide sequence as defined by a specific SEQ ID number, or over a shorter length, such as over the length of a fragment taken from a larger defined polypeptide sequence, such as a fragment of at least 15, at least 20, at least 30, at least 40, at least 50, at least 70, or at least 150 consecutive residues. Such lengths are merely exemplary, and it should be understood that any segment length supported by the sequences shown herein in tables, figures, or sequence lists can be used to describe lengths on which percentage identity can be measured.

[0090] As used herein, “reproducibility” of an XTEN polypeptide amino acid sequence refers to trimer reproducibility and can be measured by a computer program or algorithm or by other means known in the art. Trimer reproducibility of an XTEN polypeptide amino acid sequence can be evaluated by determining the number of occurrences of overlapping trimer sequences within the polypeptide. For example, a polypeptide having 200 amino acid residues has 198 overlapping trimer sequences (trimers), but the number of unique trimer sequences depends on the amount of reproducibility within the sequence. A score reflecting the degree of trimer reproducibility throughout the polypeptide sequence (hereinafter, “subsequence score”) can be generated. In the context of this invention, “subsequence score” means the sum of the occurrences of each unique trimer construct across 200 consecutive amino acids of the polypeptide divided by the absolute number of unique trimer subsequences within the 200 amino acid sequence. Examples of such subsequence scores derived from the first 200 amino acids of repeating and non-reproducing polypeptides are presented in Example 73 of International Patent Application Publication No. WO 2010 / 091122 A1, which is incorporated herein by reference in its entirety. In some embodiments, the present invention provides BPXTEN polypeptides, each comprising at least one XTEN polypeptide, wherein the amino acid sequence of the XTEN polypeptide may have a subsequence score of less than 16, or less than 14, or less than 12, or more preferably less than 10.

[0091] As used herein, the term “substantially non-repeating XTEN polypeptide amino acid sequence” refers to an XTEN polypeptide in which there are few or no instances of four consecutive amino acids of the same amino acid type in the XTEN polypeptide amino acid sequence, and wherein the XTEN polypeptide amino acid sequence has a subsequence score of 12, or 10 or lower (as defined in the preceding paragraph), or does not exhibit a pattern of sequence motifs constituting the polypeptide sequence in order from the N-terminus to the C-terminus.

[0092] As described herein, the term "non-overlapping sequence motif" includes completely non-overlapping sequence motifs as well as partially non-overlapping sequence motifs, provided that the partially non-overlapping sequence motifs are not completely overlapping.

[0093] A “vector” is a nucleic acid molecule that preferably replicates itself in a suitable host, transferring an inserted nucleic acid molecule into and / or between host cells. This term includes vectors that primarily function to insert DNA or RNA into cells, replication vectors that primarily function to replicate DNA or RNA, and expression vectors that function to transcribe and / or translate DNA or RNA. It also includes vectors that provide more than one of the above functions. An “expression vector” is a polynucleotide that, when introduced into a suitable host cell, can be transcribed and translated into a polypeptide. An “expression system” generally means a suitable host cell containing an expression vector that can function to produce the desired expression product.

[0094] As used in this article, the term "t" 1 / 2 "This means that the calculation is ln(2) / K" el The terminal half-life of K. el The terminal elimination rate constant is calculated through linear regression of the terminal linear portion of the logarithmic concentration versus time curve. Half-life generally refers to the time required for half the amount of applied substance deposited in a living organism to be metabolized or eliminated by normal biological processes. The term "t" is used in this context. 1 / 2 The terms “terminal half-life”, “elimination half-life”, and “cyclic half-life” can be used interchangeably in this article.

[0095] The terms “antigen,” “target antigen,” or “immunogen” are used interchangeably in this document to refer to an antibody fragment or a therapeutic agent based on the antibody fragment that binds to or has a structure or binding determinant specific to it.

[0096] As used herein, the term "payload" refers to a protein or peptide sequence having biological or therapeutic activity; a pair of pharmacophores for small molecules. Examples of payloads include, but are not limited to, cytokines, enzymes, hormones, and blood factors and growth factors. Payloads may also contain genetically fused or chemically conjugated portions, such as chemotherapeutic agents, antiviral compounds, toxins, or contrast agents. These conjugated portions may be linked to the remainder of the polypeptide via a linker, which may be cleavable or non-cleavable.

[0097] As used herein, the terms “treatment” or “treating,” “relief,” and “improvement” are used interchangeably and refer to methods for obtaining or achieving beneficial or desired results, including but not limited to therapeutic and / or preventive benefits. “Therapeutic benefit” means the eradication or improvement of an underlying condition that is being treated. Additionally, a therapeutic benefit is achieved by eradicating or improving one or more physiological symptoms associated with an underlying disease condition, wherein improvement is observed in a person or animal, although the person or animal may still suffer from the underlying condition. For preventive benefits, the composition may be administered to a person or animal at risk of developing a specific disease condition, or to a person or animal reporting one or more physiological symptoms of a disease, even if a diagnosis of the disease cannot yet be made.

[0098] As used herein, "therapeutic effect" refers to a physiological effect, including but not limited to the cure, mitigation, improvement, or prevention of disease conditions in humans or other animals caused by the fusion polypeptide of the present invention, or otherwise enhancing the physical or mental health of humans or animals, rather than the ability to induce the production of antibodies against antigenic epitopes possessed by biologically active proteins. In particular, the determination of a therapeutically effective amount is entirely within the capabilities of those skilled in the art, as provided in the detailed disclosure herein.

[0099] As used herein, the terms "therapeuticly effective amount" and "therapeuticly effective dose" refer to an amount of a bioactive protein, alone or as part of a fusion protein composition, that, when administered to a human or animal at a single dose or repeated doses, is capable of having any detectable beneficial effect on any symptom, aspect, measurement parameter, or characteristic of a disease state or condition. Such an effect need not be absolutely beneficial. A disease state can refer to a symptom or disease.

[0100] As used herein, the term “therapeutic effective dosing regimen” refers to a schedule of continuous administration of a bioactive protein, alone or as part of a fusion protein composition, wherein the dose is given in a therapeutically effective amount to result in a sustained beneficial effect on any symptom, aspect, measurement parameter, or characteristic of a disease state or condition.

[0101] Fusion Peptides

[0102] This document discloses polypeptides comprising one or more extended recombinant polypeptides (XTENs or XTENs) (described more fully below) that may be fused or otherwise conjugated to another polypeptide, particularly a bioactive polypeptide, wherein the embodiments described herein are referred to as BPXTENs.

[0103] In some embodiments, the polypeptide comprises a first XTEN polypeptide (e.g., those described hereinafter in the “Extended Recombinant Polypeptides (XTEN)” section or elsewhere herein). In some embodiments, the polypeptide also comprises a second XTEN polypeptide (e.g., those described hereinafter in the “Extended Recombinant Polypeptides (XTEN)” section or elsewhere herein). In some embodiments, the polypeptide comprises an XTEN polypeptide at or near its N-terminus (“N-terminal XTEN”). In some embodiments, the polypeptide comprises an XTEN polypeptide at or near its C-terminus (“C-terminal XTEN”). In some embodiments, the polypeptide comprises both an N-terminal XTEN polypeptide and a C-terminal XTEN polypeptide. In some embodiments, the first XTEN polypeptide is an N-terminal XTEN polypeptide, and the second XTEN polypeptide is a C-terminal XTEN polypeptide.

[0104] The polypeptide may also contain a bioactive polypeptide (“BP”) linked to an XTEN polypeptide, thereby forming an XTEN-containing fusion polypeptide referred to herein as the “BPXTEN” polypeptide.

[0105] The XTEN polypeptide may comprise one or more barcode fragments (configured for release) that can be released from the XTEN polypeptide after the fusion polypeptide (or BPXTEN) has been digested by a protease (described more fully below). In some embodiments, each barcode fragment is different in sequence and molecular weight from all other peptide fragments (including all other barcode fragments, if present) that can be released from the polypeptide after complete digestion by the protease.

[0106] (Fusion) polypeptides may include, for example, one or more reference fragments (described more fully below) that can be released from the polypeptide after protease digestion, the protease digestion releasing barcode fragments from the polypeptide. In some embodiments, each reference fragment may be a single reference fragment that is different in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide after protease digestion.

[0107] Extended recombinant peptide (XTEN)

[0108] Chain length and amino acid composition

[0109] In some embodiments, the XTEN polypeptide comprises at least 150 amino acids. In some embodiments, the XTEN polypeptide has a length of 150 to 3,000 amino acids, or a length of 150 to 1,000 amino acids, or a length of at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 amino acids. In some embodiments, at least 90% of the amino acid residues of the XTEN polypeptide are glycine (G), alanine (A), serine (S), threonine (T), glutamate (E), or proline (P). In some embodiments, at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid residues of the XTEN polypeptide are selected from G, A, S, T, E, or P. In some embodiments, the XTEN polypeptide comprises at least four different types of G, A, S, T, E, or P amino acids. In some embodiments, the XTEN polypeptide is characterized in that it comprises at least 150 amino acids; at least 90% of the amino acid residues of the XTEN polypeptide are G, A, S, T, E, or P, and it comprises at least four different types of amino acids selected from G, A, S, T, E, and P, with any other non-overlapping sequence motifs constituting the XTEN polypeptide being substantially randomized. In some embodiments, the fusion polypeptide containing XTEN (e.g., a fusion polypeptide comprising a bioactive polypeptide conjugated thereto) comprises a first XTEN polypeptide and a second XTEN polypeptide. In some embodiments, the sum of the total number of amino acids in the first XTEN and the total number of amino acids in the second XTEN polypeptide is at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, or at least 800 amino acids.

[0110] Non-overlapping sequence motifs

[0111] In some embodiments, the XTEN peptides provided herein comprise or are formed by a plurality of non-overlapping sequence motifs. In some embodiments, at least one non-overlapping sequence motif is reproducible (or repeated at least twice in the XTEN), and at least one other non-overlapping sequence motif is non-reproducible (or found only once in the XTEN). In some embodiments, the plurality of non-overlapping sequence motifs comprise a set of (reproducible) non-overlapping sequence motifs, each of which is repeated at least twice in the XTEN; and a non-overlapping (non-reproducible) sequence motif that appears (or is found) only once in the XTEN. In some embodiments, each non-overlapping sequence motif is 9 to 14 (or 10 to 14, or 11 to 13) amino acids in length. In some embodiments, each non-overlapping sequence motif is 12 amino acids in length. In some embodiments, the plurality of non-overlapping sequence motifs comprise a set of non-overlapping (reproducible) sequence motifs, each of which is repeated at least twice in the XTEN; and is 9 to 14 amino acids in length. In some embodiments, the set of (reproduced) non-overlapping sequence motifs comprises the 12-mer sequence motifs identified herein by SEQ ID NO: 182-203 and 1715-1722 in Table 1. In some embodiments, the set of (reproduced) non-overlapping sequence motifs comprises the 12-mer sequence motifs identified herein by SEQ ID NO: 186-189 in Table 1. In some embodiments, the set of (reproduced) non-overlapping sequence motifs comprises at least two, at least three, or all four of the 12-mer sequence motifs of SEQ ID NO: 186-189 in Table 1.

[0112] Table 1. Exemplary 12-merchant sequence motifs used to construct XTEN

[0113]

[0114]

[0115] * indicates an individual motif sequence, which, when used together in various permutations, results in a "family sequence".

[0116] Barcode fragment

[0117] In some embodiments, the polypeptides provided herein comprise a barcode fragment (e.g., a first, second, or third barcode fragment of an XTEN polypeptide) that can be released from the polypeptide upon digestion by a protease. In some embodiments, the barcode fragment is part of an XTEN that includes at least a portion of a (non-reproducible, non-overlapping) sequence motif that appears (or is found) only once within the XTEN; and is different in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide upon complete digestion by a protease. Those skilled in the art will understand that the term "barcode fragment" (or "barcode" or "barcode sequence") can refer to a portion of an XTEN identified herein that is cleavably fused within the polypeptide, or a peptide fragment obtained from the polypeptide.

[0118] In some embodiments, the barcode fragment does not include an N-terminal or C-terminal amino acid of the XTEN peptide. As described more fully below or anywhere herein, in some embodiments, the barcode fragment is releasable (configured to be releasable) after Glu-C digestion of the fusion peptide. In some embodiments, the barcode fragment does not include a glutamate immediately adjacent to another glutamate in the XTEN peptide. In some embodiments, the barcode fragment has a glutamate at its C-terminus. Those skilled in the art will understand that when cleavably fused within the XTEN peptide, the C-terminus of the barcode fragment may refer to the “last” (or C-terminal) amino acid residue within the barcode fragment, even if other “non-barcode” amino acid residues are located at the C-terminus of the barcode fragment within the same XTEN peptide. In some embodiments, the barcode fragment has an N-terminal amino acid immediately preceding a glutamate residue. In some embodiments, the glutamate residue preceding the N-terminal amino acid is not immediately adjacent to another glutamate residue. In some embodiments, the barcode fragment does not include a glutamate residue at a position other than the C-terminus of the barcode fragment, unless the glutamate is immediately followed by proline. In some embodiments, the barcode fragment is located 10 to 150, or 10 to 125 amino acids from the N-terminus or C-terminus of the polypeptide. In some embodiments, the barcode fragment is located within or at the location of the N-terminus of the polypeptide, or at a location within a range of 300, 280, 260, 250, 240, 220, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 48, 40, 36, 30, 24, 20, 12, or 10 amino acids from or at the location of the N-terminus, or within any of the foregoing. In some embodiments, the barcode fragment is located within 200, 150, 100, or 50 amino acids from the N-terminus of the polypeptide. In some embodiments, the barcode fragment is located between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the N-terminus of the polypeptide. In some embodiments, the barcode fragment is located within 300, 280, 260, 250, 240, 220, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 48, 40, 36, 30, 24, 20, 12, or 10 amino acids from the C-terminus of the polypeptide, or within any of the foregoing. In some embodiments, the barcode fragment is located within 200, 150, 100, or 50 amino acids from the C-terminus of the polypeptide. In some embodiments, the barcode fragment is located between 10 and 200, 30 and 200, 40 and 150, or 50 and 100 amino acids from the C-terminus of the polypeptide.In some embodiments, the barcode fragment does not include the N-terminal or C-terminal amino acid of the peptide; does not include glutamic acid immediately adjacent to another glutamic acid in the XTEN; has glutamic acid at its C-terminus; has an N-terminal amino acid immediately preceding a glutamic acid residue; and (v) is positioned 10 to 150, or 10 to 125 amino acids away from the N-terminus or C-terminus of the peptide. In some embodiments, the glutamic acid residue preceding the N-terminal amino acid is not immediately adjacent to another glutamic acid residue. In some embodiments, the barcode fragment does not include glutamic acid residues at positions other than the C-terminus of the barcode fragment, unless the glutamic acid is immediately followed by proline. In some embodiments, for a barcoded XTEN peptide fused to a bioactive peptide, at least one barcode fragment (or at least two, or three barcode fragments) contained in the barcoded XTEN is positioned at least 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, or 300 amino acids away from the bioactive peptide. In some embodiments, the length of the barcode segment is at least 4, at least 5, at least 6, at least 7, or at least 8 amino acids. In some embodiments, the length of the barcode segment is at least 4 amino acids. In some embodiments, the length of the barcode segment is 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids, or within any of the aforementioned values. In some embodiments, the length of the barcode segment is 4 to 20, 5 to 15, 6 to 12, or 7 to 10 amino acids. In some embodiments, the barcode segment is selected from SEQ ID NO: 8020-8030 (BAR001-BAR011) in Table 2.

[0119] Table 2. Exemplary barcode fragments that can be released after Glu-C digestion

[0120] amino acid sequence SEQ ID NO: SPATSGSTPE BAR001 8020 GSAPATSE BAR002 8021 GSAPGTATE BAR003 8022 GSAPGTE BAR004 8023 PATSGPTE BAR005 8024 SASPE BAR006 8025 PATSGSTE BAR007 8026 GSAPGTSAE BAR008 8027 SATSGSE BAR009 8028 SGPGSTPAE BAR010 8029 SGSE BAR011 8030

[0121] In some embodiments, the barcoded XTEN peptide contains only one barcode fragment. In some embodiments, the barcoded XTEN peptide contains a set of barcode fragments, including a first barcode fragment, such as those described above or anywhere else herein. In these embodiments, each member of this set of barcode sequences may be distinguished from all other barcode sequences based on amino acid sequence or molecular weight (wherein these methods for distinguishing different barcode sequences will be relevant). In some embodiments, the set of barcode fragments includes a second barcode fragment (or further barcode fragments), such as those described above or anywhere else herein. In some embodiments, the set of barcode fragments includes a third barcode fragment, such as those described above or anywhere else herein. This set of barcode fragments fused within the N-terminal XTEN peptide may be referred to as an N-terminal group barcode (“N-terminal group”). This set of barcode fragments fused within the C-terminal XTEN peptide may be referred to as a C-terminal group barcode (“C-terminal group”). In some embodiments, the N-terminal group includes a first barcode segment and a second barcode segment. In some embodiments, the N-terminal group further includes a third barcode segment. In some embodiments, the C-terminal group includes a first barcode segment and a second barcode segment. In some embodiments, the C-terminal group further includes a third barcode segment. In some embodiments, a second barcode segment is positioned at the N-terminus of the first barcode segment in the same group. In some embodiments, a second barcode segment is positioned at the C-terminus of the first barcode segment in the same group. In some embodiments, a third barcode segment is positioned at the N-terminus of both the first and second barcode segments. In some embodiments, a third barcode segment is positioned at the C-terminus of both the first and second barcode segments. In some embodiments, a third barcode segment is positioned between the first and second barcode segments. In some embodiments, the polypeptide comprises a set of barcode fragments, including a first barcode fragment, a further (second) barcode fragment, and at least one additional barcode fragment, wherein each barcode fragment in the set is part of a second XTEN polypeptide and is different in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide after complete digestion by the protease.

[0122] Example of XTEN with barcode

[0123] Table 3a shows 13 exemplary barcoded XTEN amino acid sequences containing one barcode (e.g., SEQ ID NO: 8002-8003, 8005-8009, and 8013), two barcodes (e.g., SEQ ID NO: 8001, 8004, 8010, and 8012), or three barcodes (e.g., SEQ ID NO: 8011). Of these 13 exemplary barcoded XTEN peptides, six (SEQ ID NO: 8001-8003, 8008-8009, and 8011) can be fused to the C-terminus of a bioactive protein, and seven (SEQ ID NO: 8004-8007, 8010, and 8012-8013) can be fused to the N-terminus of a bioactive protein. In some embodiments, the XTEN polypeptide has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with sequences selected from SEQ ID NO:8001-8019 in Table 3a.

[0124] Table 3a. Exemplary XTEN with barcodes

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140] In some embodiments, a barcoded XTEN peptide can be obtained by preparing one or more mutations in a generic XTEN peptide, such as any peptide listed in Table 3b, according to one or more of the following criteria: minimizing sequence changes in the XTEN peptide, minimizing changes in the amino acid composition of the XTEN peptide, substantially maintaining the net charge of the XTEN peptide, substantially maintaining (or improving) the low immunogenicity of the XTEN peptide, and substantially maintaining (or improving) the pharmacokinetic properties of the XTEN peptide. In some embodiments, the amino acid sequence of the XTEN peptide has at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with any of the SEQ ID NOs: 676-734 listed in Table 3b. In some embodiments, an XTEN sequence having at least 90% (e.g., at least 92%, at least 95%, at least 98%, or at least 99%) but less than 100% sequence identity with any of the SEQ ID NO: 676-734 listed in Table 3b is obtained by one or more mutations (e.g., fewer than 10, fewer than 8, fewer than 6, fewer than 5, fewer than 4, fewer than 3, or fewer than 2 mutations) of the corresponding sequences from Table 3b. In some embodiments, one or more mutations comprise the deletion, insertion, substitution, or replacement of glutamate residues, or any combination thereof. In some embodiments, wherein the amino acid sequence of the XTEN polypeptide differs from any of the SEQ ID NOs: 676-734 listed in Table 3b, but has at least 90% (e.g., at least 92%, at least 95%, at least 98%, or at least 99%) sequence identity, at least 80%, at least 90%, at least 95%, at least 97%, or about 100% of the difference between the XTEN polypeptide amino acid sequence and the corresponding sequence in Table 3b involves the deletion, insertion, substitution, or replacement of a glutamate residue, or any combination thereof. In some such embodiments, at least 80%, at least 90%, at least 95%, at least 97%, or about 100% of the difference between the XTEN polypeptide amino acid sequence and the corresponding sequence in Table 3b involves the substitution, or replacement of a glutamate residue, or both. As used herein, the term “substitution of the first amino acid” refers to the replacement of a second amino acid residue with a first amino acid residue, resulting in the second amino acid residue appearing at the substitution position in the obtained sequence. For example, “substitution of glutamic acid” refers to replacing a non-glutamic acid residue (e.g., serine (S)) with a glutamic acid (E) residue. As used herein, the term “substitution of the first amino acid” refers to replacing a first amino acid residue with a second amino acid residue, resulting in the first amino acid residue appearing at the substitution position in the obtained sequence. For example, “substitution of glutamic acid” refers to replacing a glutamic acid residue with a non-glutamic acid residue (e.g., serine (S)).

[0141] Table 3b. Exemplary generic XTENs for conversion into barcode-enabled XTENs

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165] In some embodiments, in order to construct the sequence of the barcoded XTEN peptide, amino acid mutations are performed on the XTEN peptides of medium length among those in Table 3b and the XTEN peptides of longer length than those in Table 3b, such as those XTEN peptides in which one or more 12-mer motifs of Table 1 are added to the N-terminus or C-terminus of the universal XTEN in Table 3b.

[0166] Further examples of the generic XTEN polypeptide amino acid sequences that can be used in accordance with this disclosure are disclosed in U.S. Patent Publications 2010 / 0239554 A1, 2010 / 0323956 A1, 2011 / 0046060 A1, 2011 / 0046061 A1, 2011 / 0077199 A1 or 2011 / 0172146 A1, or International Patent Publications WO 2010091122 A1, WO 2010144502 A2, WO 2010144508 A1, WO 2011028228 A1, WO 2011028229 A1, WO 2011028344 A2, WO2014 / 011819 A2 or WO In 2015 / 023891, the disclosures of the aforementioned patents are each expressly incorporated herein by reference.

[0167] In some embodiments, a barcoded XTEN polypeptide (“N-terminal XTEN”) fused within the polypeptide chain adjacent to the N-terminus of the polypeptide chain may be attached to a His tag containing multiple His residues, including six to eight His residues at the N-terminus, to facilitate the purification of the fusion polypeptide. In some embodiments, a barcoded XTEN polypeptide (“C-terminal XTEN polypeptide”) fused within the polypeptide chain at the C-terminus of the polypeptide chain may contain or be attached to a sequence EPEA at the C-terminus to facilitate the purification of the fusion polypeptide. In some embodiments, the fusion peptide comprises both an N-terminal barcoded XTEN peptide and a C-terminal barcoded XTEN peptide, wherein the N-terminal barcoded XTEN is attached to a His tag comprising a plurality of poly(His) residues, including six to eight His residues at the N-terminus; and wherein the C-terminal barcoded XTEN peptide is attached to a sequence EPEA at the C-terminus, thereby facilitating the purification of the fusion peptide by chromatographic methods known in the art, for example to a purity of at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, including but not limited to IMAC chromatography, C-tagXL affinity matrix, and other such methods, including but not limited to those described in the example sections below.

[0168] Protease digestion

[0169] As described above or anywhere else herein, the barcode fragment can be cleavably fused within the XTEN polypeptide and can be released from the XTEN polypeptide (configured for release) upon digestion by a protease. In some embodiments, the protease is a Glu-C protease. In some embodiments, the protease cleaves at the C-terminal side of a glutamate residue, which is subsequently not a proline. Those skilled in the art will understand that barcoded XTEN polypeptides (XTEN polypeptides containing the barcode fragment therein) are designed to achieve high efficiency, precision, and accuracy of protease digestion. For example, those skilled in the art will understand that adjacent Glu-Glu (EE) residues in the XTEN sequence can result in various cleavage patterns after Glu-C digestion. Accordingly, when the Glu-C protease is used for barcode release, the barcoded XTEN polypeptide or barcode fragment may not contain any Glu-Glu (EE) sequence. Those skilled in the art will also understand that if present in the fusion polypeptide, the dipeptide Glu-Pro (EP) sequence may be uncleavable by the Glu-C protease during the barcode release process.

[0170] BPXTEN structural configuration

[0171] In some embodiments, the BPXTEN fusion protein comprises a single BP polypeptide and a single XTEN polypeptide. Such BPXTEN proteins may have at least the following conformations, each listed in N-terminal to C-terminal orientation: BP-XTEN; XTEN-BP; BP-S-XTEN; and XTEN-S-BP, where “S” is a spacer sequence as described below.

[0172] In some embodiments, the BPXTEN protein comprises a C-terminal XTEN polypeptide and optionally a spacer region sequence (S) between the XTEN polypeptide and the BP polypeptide. Such a BPXTEN protein can be represented by Formula I (described as N-terminus to C-terminus):

[0173] (BP)-(S) x -(XTEN) (I),

[0174] Wherein BP is a bioactive protein as described below; S is a spacer sequence having 1 to 50 amino acid residues, which may optionally include a BP release segment (as described more fully below); x is 0 or 1; and XTEN may be any XTEN polypeptide described herein.

[0175] In some embodiments, the BPXTEN protein comprises an N-terminal XTEN polypeptide and optionally a spacer region sequence (S) between the XTEN polypeptide and the BP protein. Such a BPXTEN protein can be represented by Formula II (described as N-terminus to C-terminus):

[0176] (XTEN)-(S) x -(BP) (II),

[0177] Wherein BP is a bioactive protein as described below; S is a spacer sequence having 1 to 50 amino acid residues, which may optionally include a BP release segment (as described more fully below); x is 0 or 1; and XTEN may be any XTEN polypeptide as described herein.

[0178] In some embodiments, the BPXTEN protein comprises both an N-terminal XTEN polypeptide and a C-terminal XTEN polypeptide. Such BPXTEN proteins (e.g., Figure 1-2 XPAT in the equation can be represented by Equation III:

[0179] (XTEN)-(S) y -(BP)-(S) z -(XTEN) (III)

[0180] Wherein BP is a bioactive protein as described below; S is a spacer sequence having 1 to 50 amino acid residues, which may optionally include a BP release segment (as described more fully below); y is 0 or 1; z is 0 or 1; and XTEN may be any XTEN polypeptide as described herein.

[0181] Bioactive peptides

[0182] Bioactive proteins (BPs) that can be fused to one or more XTEN polypeptides (as described herein), particularly those disclosed below, are contained in the sequences identified herein by means of Tables 4a-4h and 6a-6f, together with their corresponding nucleic acid and amino acid sequences, which are well known in the art. Descriptions and sequences of these BPs are available in public databases such as Chemical Abstracts Services Databases (e.g., CAS Registry), GenBank, The Universal Protein Resource (UniProt), and subscription-based databases such as GenSeq (e.g., Derwent). The polynucleotide sequence encoding the BP can be a wild-type polynucleotide sequence encoding a native BP (e.g., full-length or mature), or in some cases, a variant of a wild-type polynucleotide sequence (e.g., a polynucleotide encoding a wild-type, bioactive protein) whose nucleotide sequence has been optimized, for example, for expression in a particular species; or a polynucleotide encoding a variant of a wild-type protein, such as a site-directed mutant or allelic variant. Using methods known in the art and / or in conjunction with the guidance and methods provided herein, the BPXTEN constructs considered in this invention can be generated using wild-type or common cDNA sequences or codon-optimized variants of BP, entirely within the capabilities of those skilled in the art.

[0183] The BP used in BPXTEN proteins disclosed herein (e.g., fusion peptides comprising at least one BP and at least one XTEN peptide) may comprise any protein having a biological, therapeutic, preventive, or diagnostic purpose or function, or, when administered to humans or animals, may be used to mediate biological activity or to prevent or improve a disease, symptom, or condition. Particularly advantageous are BPs for which increased pharmacokinetic parameters, increased solubility, increased stability, activity masking, or some other enhanced pharmaceutical property are sought, or for those BPs for which increased terminal half-life would improve efficacy, safety, or result in reduced dosing frequency and / or improved patient compliance. Thus, various objectives can be considered in preparing BPXTEN fusion protein compositions, including improving the therapeutic efficacy of the biologically active compound compared to BPs not linked to an XTEN peptide, by, for example, increasing in vivo exposure or the length of time BPXTEN remains within the therapeutic window when administered to humans or animals.

[0184] BP can be a natural full-length protein, or it can be a fragment or sequence variant of a bioactive protein that retains at least a portion of the biological activity of the natural protein.

[0185] In one embodiment, the BP incorporated into a human or animal composition may be a recombinant polypeptide having a sequence corresponding to a protein found in nature. In another embodiment, the BP may be a sequence variant, fragment, homolog, or mimic of the natural sequence, retaining at least a portion of the biological activity of the natural BP. In a non-limiting example, the BP may be a sequence showing at least about 80% sequence identity with a protein sequence selected from Tables 4a-4h, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity. In a further non-limiting example, BP may be a bispecific sequence comprising a first binding domain and a second binding domain, wherein the first binding domain has a specific binding affinity for tumor-specific markers or antigens of target cells, and shows at least about 80% sequence identity with the paired VL and VH sequences of the anti-CD3 antibodies identified in Table 6f, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 9 6%, 97%, 98%, 99%, or 100% sequence identity; and wherein the second binding domain, which has specific binding affinity for effector cells, shows at least about 80% sequence identity with the paired VL and VH sequences of the anti-target cell antibodies identified in Table 6a, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity. In one embodiment, the BPXTEN fusion protein may comprise a single BP protein linked to an XTEN polypeptide. In another embodiment, the BPXTEN protein may comprise a first BP and a second molecule of the same BP, resulting in a fusion protein comprising two BPs (e.g., two glucagon molecules or two hGH molecules) linked to one or more XTEN polypeptides.

[0186] Generally, when used in vivo or utilized in in vitro assays, a BP exhibits binding specificity to a given target (or a given number of targets) or another desired biological property. For example, a BP can be an agonist, receptor, ligand, antagonist, enzyme, antibody (e.g., monospecific or bispecific), or hormone. Of particular interest are BPs intended for or known to be applicable to diseases or conditions, wherein the natural BP has a relatively short terminal half-life, and enhancement of its pharmacokinetic parameters (which may optionally be released from the fusion protein by cleavage of the spacer region sequence) would allow for less frequent dosing or enhanced pharmacological effects. Also of interest are BPs having a narrow therapeutic window between the minimum effective dose or blood concentration (Cmin) and the maximum tolerated dose or blood concentration (Cmax). In this case, the linking of a BP to a fusion protein containing a selected XTEN peptide sequence can lead to improvements in these properties compared to a BP not linked to one or more XTEN peptides, making it more suitable for use as a therapeutic or prophylactic agent.

[0187] Glucose-regulated peptides

[0188] Endocrine and obesity-related diseases or conditions have reached epidemic proportions in most developed countries and represent a huge and ever-increasing healthcare burden, encompassing a wide variety of conditions affecting the body's organs, tissues, and circulatory systems. Of particular concern are endocrine and obesity-related diseases and conditions, primarily diabetes, one of the leading causes of death in the United States.

[0189] Most metabolic processes in glucose homeostasis and insulin response are regulated by multiple peptides and hormones, and many of these peptides and hormones, as well as their analogues, have been found to be effective in the treatment of metabolic diseases and conditions. Many of these peptides tend to be highly homologous to each other, even when they have opposing biological functions. Peptides that increase glucose include the peptide hormone glucagon, while those that decrease glucose include exendin-4, glucagon-like peptide-1, and amylin. However, even when intensified with the use of small molecule drugs, the use of therapeutic peptides and / or hormones has achieved limited success in the management of these diseases and conditions. In particular, dosage optimization is important for drugs and biologics used to treat metabolic diseases, especially those with narrow therapeutic windows. Hormones in general and peptides involved in glucose homeostasis often have narrow therapeutic windows. This narrow therapeutic window, coupled with the fact that such hormones and peptides typically have short half-lives (requiring frequent dosing to achieve clinical benefit), leads to difficulties in the management of these patients. While chemical modifications to therapeutic proteins, such as PEGylation, can alter their in vivo clearance and subsequent serum half-life, this requires additional manufacturing steps and results in heterogeneous end products. Furthermore, unacceptable side effects from prolonged administration have been reported. Alternatively, genetic modifications via the fusion of the Fc domain with the therapeutic protein or peptide can increase the size of the therapeutic protein, reduce clearance by the kidneys, and promote lysosomal recycling via the FcRn receptor. Unfortunately, the Fc domain cannot fold efficiently during recombinant expression and tends to form insoluble precipitates called inclusion bodies. These inclusion bodies must be dissolved, and the functional protein must be reactivated; this is a time-consuming, inefficient, and costly process.

[0190] Therefore, one aspect of the present invention is to incorporate peptides relating to glucose homeostasis, insulin resistance, and obesity (collectively, “glucose-regulating peptides”) into a BPXTEN fusion protein to produce compositions effective in the treatment of glucose, insulin, and obesity disorders, diseases, and related conditions. Suitable glucose-regulating peptides that can be linked to the XTEN polypeptides disclosed herein to produce BPXTEN proteins (which in particular include all bioactive peptides) increase glucose-dependent insulin secretion through pancreatic β-cells or enhance the action of insulin. Glucose-regulating peptides may also include bioactive peptides that stimulate transcription of the proinsulin gene in pancreatic β-cells. Furthermore, glucose-regulating peptides may also include bioactive peptides that slow gastric emptying time and reduce food intake. Glucose-regulating peptides may also include bioactive peptides that inhibit glucagon release from α-cells of Langerhans islands. Table 4a provides a non-limiting list of glucose-regulating peptide sequences that may be covered by the BPXTEN fusion protein of the present invention. The glucose-regulating peptide of the BPXTEN composition of the present invention disclosed herein may be a peptide that shows at least about 80% sequence identity with the amino acid sequence selected from Table 4a (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity).

[0191] Table 4a: Glucose-regulating peptides

[0192]

[0193]

[0194]

[0195]

[0196] "Adrenal medullaris" or "ADM" refers to human adrenal medullaris peptide hormone and species and sequence variants of ADM that possess at least a portion of the biological activity of mature ADM. ADM is generated from a 185-amino acid prohormone through sequential enzymatic cleavage and amidation, resulting in a 52-amino acid bioactive peptide with a measured plasma half-life of 22 minutes. The ADM-containing fusion protein of the present invention can be used specifically in diabetes for stimulating insulin secretion from pancreatic islet cells for glucose regulation, or in humans or animals with persistent hypotension. The complete genomic basis of human AM has been reported (Ishimitsu et al., 1994, Biochem. Biophys. Res. Commun 203:631-639), and analogues of the ADM peptide have been cloned, as described in U.S. Patent No. 6,320,022.

[0197] "Amylin" refers to human peptide hormones called amylin, pramlinide, and species variations of amylin possessing at least a portion of the biological activity of mature amylin, as described in U.S. Patent No. 5,234,906. Amylin is a 37-amino acid polypeptide hormone co-secreted by pancreatic β-cells with insulin in response to nutrient intake (Koda et al., 1992, Lancet 339:1179-1180), and has been reported to regulate several key pathways of carbohydrate metabolism, including glucose incorporation into glycogen. The amylin-containing fusion protein of the present invention can complement the action of insulin, which regulates the rate of glucose loss from circulation and its uptake by peripheral tissues. Amylin analogues have been cloned, as described in U.S. Patent Nos. 5,686,411 and 7,271,238.

[0198] Amylin mimics that retain biological activity can be produced. For example, pramlinin has the sequence KCNTATCATNRLANFLVHSSNNFGPILPPTNVGSNTY (SEQ ID NO:43), wherein an amino acid from the rat amylin sequence replaces an amino acid in the human amylin sequence. In one embodiment, the invention contemplates a fusion protein comprising an amylin mimic containing the sequence KCNTATCATX1RLANFLVHSSNNFGX2ILX2X2TNVGSNTY (SEQ ID NO:44), wherein X1 is independently N or Q and X2 is independently S, P, or G. In one embodiment, the amylin mimic incorporated into BPXTEN may have the sequence KCNTATCATNRLANFLVHSSNNFGGILGGTNVGSNTY (SEQ ID NO:45). In another embodiment, wherein the amylin mimic is used at the C-terminus of BPXTEN, the mimic may have the sequence KCNTATCATNRLANFLVHSSNNFGGILGGTNVGSNTY(NH2) (SEQ ID NO:46).

[0199] "Calcitonin" (CT) refers to human calcitonin protein and species and sequence variants of CT with at least a portion of the biological activity of the mature CT, including salmon calcitonin ("sCT"). CT is a 32-amino acid peptide cleaved from a larger prothyroid hormone, which appears to function in the nervous and vascular systems, but has also been reported as a potent hormonal mediator of the satiety reflex (reviewed in Becker, JCEM, 89(4):1512-1525 (2004) and Sexton, Current Medicinal Chemistry 6:1067-1093 (1999)). The calcitonin-containing fusion protein of the present invention can be used specifically for the treatment of osteoporosis and as a therapy for Paget's disease. Synthetic calcitonin peptides have been produced as described in U.S. Patent Nos. 5,175,146 and 5,364,840.

[0200] "Calcitonin gene-related peptide" or "CGRP" refers to human CGRP peptide and species and sequence variants of CGRP possessing at least a portion of the biological activity of mature CGRP, which is a member of the calcitonin family of peptides and exists in humans in two forms: α-CGRP (a 37-amino acid peptide) and β-CGRP. CGRP shares 43-46% sequence identity with human amylin. The CGRP-containing fusion protein of the present invention can be specifically used to reduce the incidence of diabetes-related diseases, improve hyperglycemia and insulin deficiency, inhibit lymphocyte infiltration into the islets of Langerhans, and protect β-cells from autoimmune destruction. Methods for preparing synthetic and recombinant CGRP are described in U.S. Patent No. 5,374,618.

[0201] "Cholecystokinin" or "CCK" refers to human CCK peptides and species and sequence variants of those possessing at least a portion of the biological activity of mature CCK. CCK-58 is the mature sequence, while the CCK-33 amino acid sequence, first identified in humans, is the major circulating form of this peptide. The CCK family also includes an in vivo C-terminal fragment of 8 amino acids ("CCK-8"), the pentapeptide gastrin or CCK-5 of the C-terminal peptide CCK (29-33), and the CCK-4 of the C-terminal tetrapeptide CCK (30-33). CCK is a peptide hormone of the gastrointestinal system responsible for stimulating the digestion of fats and proteins. The fusion protein of the present invention containing CCK-33 and CCK-8 can be specifically used to reduce the increase in circulating glucose and enhance the increase in circulating insulin after dietary intake. Analogs of CCK-8 have been prepared as described in U.S. Patent No. 5,631,230.

[0202] "Exendin-3" refers to a glucose-regulating peptide isolated from the beaded venom lizard (Heloderma horridum) and its sequence variants possessing at least a portion of the biological activity of mature exendin-3. Exendin-3 amides are specific exendin receptor antagonists that mediate increases in pancreatic cAMP and the release of insulin and amylase. The exendin-3-containing fusion protein of this invention can be specifically used to treat diabetes and insulin resistance. The sequence and methods for its determination are described in U.S. Patent 5,424,286.

[0203] "Exendin-4" refers to the glucose-regulating peptide found in the saliva of the American venomous lizard (Heloderma suspectum) and its species and sequence variants, including the native 39-amino acid sequence HGEGTFTSDLSKQMEEEAVRLFIEYLKNGGPSSGAPPPS (SEQ ID NO:47) and homologous sequences and peptide mimics and their variants; such as native sequences from primates and non-natural sequences possessing at least a portion of the biological activity of mature exendin-4. Exendin-4 is an incretin polypeptide hormone that lowers blood glucose, promotes insulin secretion, slows gastric emptying, and improves satiety, providing significant improvement in postprandial hyperglycemia. Table 4b shows sequences from a wide variety of species, while Table 4c shows a list of synthetic GLP-1 analogs; all of these are considered for use in the BPXTEN protein described herein.

[0204] Fibroblast growth factor 21, or "FGF-21", refers to the human protein encoded by the FGF-21 gene, or species and sequence variants thereof possessing at least a portion of the biological activity of mature FGF-21. FGF-21 stimulates glucose uptake in adipocytes, but not in other cell types; this effect is additive with respect to insulin activity. The FGF-21-containing fusion protein of the present invention can be specifically used to treat diabetes, including by inducing increased energy expenditure, fat utilization, and lipid excretion. FGF-21 has been cloned, as disclosed in U.S. Patent No. 6,716,626.

[0205] "Fibroblast growth factor 19" or "FGF-19" means the human protein encoded by the FGF-19 gene, or a species and sequence variant thereof having at least a portion of the biological activity of mature FGF-19. FGF-19 is a protein member of the fibroblast growth factor (FGF) family. FGF-19 increases hepatic expression of the leptin receptor, metabolic rate, stimulates glucose uptake in adipocytes, and leads to weight loss in obese mouse models (Fu et al., 2004, Endocrinology 145:2504-2603). The FGF-19-containing fusion protein of the present invention can be specifically used to increase metabolic rate and reverse diet and leptin-deficient diabetes. FGF-19 has been cloned and expressed as described in U.S. Patent Application No. 20020042367.

[0206] "Gastrin" refers to the human gastrin peptide, its truncated form, and species and sequence variants possessing at least a portion of the biological activity of mature gastrin. Gastrin is primarily found in three forms: gastrin-34 ("large gastrin"); gastrin-17 ("small gastrin"); and gastrin-14 ("small gastrin"), and shares sequence homology with CCK. The gastrin-containing fusion protein of the present invention can be specifically used to treat obesity and diabetes for glucose regulation. Gastrin has been synthesized as described in U.S. Patent No. 5,843,446.

[0207] "Ghrelin" refers to a human hormone that induces satiety, or species and sequence variants thereof, including natural and processed sequences of 27 or 28 amino acids and homologous sequences. Ghrelin levels increase before meals and decrease after meals, and can lead to increased food intake and increased fat mass through its action at the hypothalamic level. The ghrelin-containing fusion proteins of the present invention can be used particularly as agonists; for example, selectively stimulating the motility of the GI tract in gastrointestinal motility disorders, accelerating gastric emptying, or stimulating the release of growth hormone. For example, ghrelin analogs with sequence substitutions or truncated variants as described in U.S. Patent No. 7,385,026 can be used particularly as fusion couples with XTEN peptides as antagonists for improving glucose homeostasis, treating insulin resistance, and treating obesity. The isolation and characterization of ghrelin have been reported (Kojima et al., 1999, Nature. 402:656-660), and synthetic analogs have been prepared by peptide synthesis as described in U.S. Patent No. 6,967,237.

[0208] "Glucagon" refers to human glucagon glucose-regulating peptide, or species and sequence variants thereof, including the natural 29-amino acid sequence and homologous sequences; for example, natural sequence variants from primates; and non-natural sequence variants having at least a portion of the biological activity of mature glucagon. As used herein, the term "glucagon" also includes peptide mimics of glucagon. The glucagon-containing fusion protein of the present invention can be particularly used to increase blood glucose levels in individuals with existing liver glycogen reserves and to maintain glucose homeostasis in diabetic patients. Glucagon has been cloned, as disclosed in U.S. Patent No. 4,826,763.

[0209] "GLP-1" refers to human glucagon-like peptide-1 and its sequence variants possessing at least a portion of the biological activity of mature GLP-1. The term "GLP-1" includes human GLP-1(1-37), GLP-1(7-37), and GLP-1(7-36) amide. GLP-1 stimulates insulin secretion, but only during periods of hyperglycemia. The safety of GLP-1 compared to insulin is enhanced by this property and the observation that the amount of insulin secreted is proportional to the magnitude of hyperglycemia. The biological half-life of GLP-1(7-37)OH is only 3 to 5 minutes (US Patent No. 5,118,666). The GLP-1-containing fusion protein of the present invention can be specifically used to treat diabetes and insulin resistance for glucose regulation. GLP-1 has been cloned and derivatives prepared as described in US Patent No. 5,118,666. Non-limiting examples of GLP-1 sequences from a wide variety of species are shown in Table 4b, while Table 4c shows sequences of many synthetic GLP-1 analogs; all of these are considered for use in the BPXTEN compositions described herein.

[0210] Table 4b: Representative naturally occurring GLP-1 homologs as BP candidates

[0211]

[0212]

[0213]

[0214]

[0215] Table 4c: Representative GLP-1 synthetic analogs

[0216]

[0217]

[0218]

[0219]

[0220] The natural GLP sequence can be described by several sequence motifs presented below. The letters in parentheses represent the amino acids that are acceptable at each sequence position: {HVY}{AGISTV}{DEHQ}{AG}{ILMPSTV}{FLY}{DINST}{ADEKNST}{ADENSTV}{LMVY}{ANRSTY}{EHIKNQRST}{AHILMQVY}{LMRT}{ADEGKQS}{ADEGKNQSY}{AEIKLMQR}{AKQRSVY}{{AILMQSTV}{GKQR}{DEKLQR}{FHLVWY}{ILV}{ADEGHIKNQRST}{ADEGNRSTW}{GILVW}{AIKLMQSV}{ADGIKNQRST}{GKRSY} (SEQ ID NO:9399). In addition, synthetic analogues of GLP-1 can be used as fusion partners of XTEN peptides to produce BPXTEN proteins with bioactivity that can be used to treat glucose-related diseases.

[0221] “GLP-2” refers to human glucagon-like peptide-2 and its sequence variants that possess at least a portion of the biological activity of mature GLP-2. More specifically, GLP-2 is a 33-amino acid peptide co-secreted by enteroendocrine cells in the small and large intestines, together with GLP-1.

[0222] "Insulin-like growth factor 1" or "IGF-1" refers to the human IGF-1 protein and species and sequence variants possessing at least a portion of the biological activity of mature IGF-1. IGF-1 consists of 70 amino acids and is primarily produced by the liver as an endocrine hormone, as well as in target tissues via paracrine / autocrine pathways. The IGF-1-containing fusion protein of this invention can be specifically used to treat diabetes and insulin resistance for glucose regulation. IGF-1 has been cloned and expressed in *E. coli* and yeast, as described in U.S. Patent No. 5,324,639.

[0223] "Insulin-like growth factor 2" or "IGF-2" refers to the human IGF-2 protein and species and sequence variants that possess at least a portion of the biological activity of mature IGF-2. IGF-2 has been cloned, as described in Bell et al., 1985, Proc Natl AcadSci US A.82:6450-4.

[0224] "Insulin-associated islet cell growth factor" (INGAP) or "pancreatic β-cell growth factor" refers to human INGAP peptide and species and sequence variants of INGAP that possess at least a portion of the biological activity of mature INGAP. The INGAP-containing fusion protein of this invention can be specifically used for the treatment or prevention of diabetes and insulin resistance. INGAP has been cloned and expressed, as described in R Rafaeloff et al., 1997, J Clin Invest. 99(9):2100–2109.

[0225] "Pituitary middle lobe peptide" or "AFP-6" refers to human pituitary middle lobe peptide and species and sequence variants of it having at least a portion of the biological activity of mature pituitary middle lobe peptide. Treatment with pituitary middle lobe peptide results in a decrease in blood pressure in normal and hypertensive humans or animals, as well as inhibition of gastric emptying activity, and involves glucose homeostasis. The pituitary middle lobe peptide-containing fusion protein of the present invention can be specifically used to treat diabetes, insulin resistance, and obesity. The pituitary middle lobe peptide and variants have been cloned as described in U.S. Patent No. 6,965,013.

[0226] "Lepin" refers to naturally occurring leptin from any species, as well as the biologically active D-isoform, or fragments and sequence variants thereof. The leptin-containing fusion protein of this invention can be specifically used to treat diabetes, for glucose regulation, insulin resistance, and obesity. Leptin has been cloned as described in U.S. Patent No. 7,112,659, and leptin analogs and fragments have been cloned as described in U.S. Patent Nos. 5,521,283, 5,532,336, PCT / US96 / 22308, and PCT / US96 / 01471.

[0227] "Neurotransferase" refers to the neurotransferase family of peptides, including neurotransferase U and S peptides, and their sequence variants. The neurotransferase U family includes various truncated or spliced ​​variants, such as FLFHYSKTQKLGKSNVVEELQSPFASQSRGYFLFRPRN (SEQ ID NO:180). An example of the neurotransferase S family is human neurotransferase S having the sequence ILQRGSGTAAVDFTKKDHTATWGRPFFLFRPRN (SEQ ID NO:181), particularly its amide form. The neurotransferase fusion protein of the present invention can be particularly used to treat obesity, diabetes, reduced food intake, and other related conditions and disorders as described herein.

[0228] "Gasitin" or "OXM" refers to human gastrin and species and sequence variants of it that possess at least a portion of the biological activity of mature OXM. OXM is a 37-amino acid peptide produced in the colon, containing the 29-amino acid sequence of glucagon, followed by an 8-amino acid carboxyl-terminal extension. The OXM-containing fusion protein of the present invention can be specifically used to treat diabetes for glucose regulation, insulin resistance, obesity, and can also be used as a weight loss therapy.

[0229] "PYY" refers to the human peptide YY polypeptide and species and sequence variants of it that possess at least a portion of the biological activity of mature PYY. The PPY-containing fusion protein of this invention can be specifically used to treat diabetes for glucose regulation, insulin resistance, and obesity. Analogs of PYY have been prepared as described in U.S. Patent Nos. 5,604,203, 5,574,010, and 7,166,575.

[0230] "Urocorticin" refers to human urocorticin peptide hormone and its sequence variants having at least a portion of the biological activity of mature urocorticin. Three human urocorticins exist: Ucn-1, Ucn-2, and Ucn-3. Further urocorticins and analogues are described in U.S. Patent No. 6,214,797. The BPXTEN protein containing urocorticin of the present invention can also be used specifically to treat or prevent conditions associated with stimulation of ACTH release, hypertension due to vasodilatory effects, inflammation mediated by factors other than ACTH elevation, high fever, appetite disorders, congestive heart failure, stress, anxiety, and psoriasis. Urocorticin-containing fusion proteins can also be combined with natriuretic peptide modules, amylin family and exendin family or GLP1 family modules to provide enhanced cardiovascular benefits, such as treating CHF by providing a beneficial vasodilatory effect.

[0231] Metabolic diseases and cardiovascular proteins

[0232] Metabolic and cardiovascular diseases represent a significant healthcare burden in most developed countries, with cardiovascular disease remaining the leading cause of death and disability in the United States and most European countries. Metabolic diseases and conditions encompass a wide range of conditions affecting the body's organs, tissues, and circulatory system.

[0233] Dyslipidemia is frequently observed in people with diabetes and in animals with cardiovascular disease; it is typically characterized by parameters such as elevated plasma triglycerides, low HDL (high-density lipoprotein) cholesterol, normal to elevated LDL (low-density lipoprotein) cholesterol, and increased levels of small, dense LDL particles in the blood. Dyslipidemia and hypertension are major contributors to increased morbidity of coronary events, kidney disease, and death in people and animals with metabolic diseases such as diabetes and cardiovascular disease.

[0234] Cardiovascular diseases can manifest as a wide range of conditions, symptoms, and changes in clinical parameters involving the heart, blood vessels, and organ systems throughout the body. These include aneurysms, angina pectoris, atherosclerosis, stroke, cerebrovascular disease, congestive heart failure, coronary artery disease, myocardial infarction, decreased cardiac output, peripheral vascular disease, hypertension, hypotension, and changes in blood markers (such as C-reactive protein, BNP, and enzymes such as CPK, LDH, SGPT, and SGOT).

[0235] Most metabolic processes and many cardiovascular parameters are regulated by multiple peptides and hormones (“metabolic proteins”), and many of these peptides and hormones, as well as their analogues, have been found to be effective in the treatment of these diseases and conditions. However, even when enhanced with the use of small molecule drugs, the use of therapeutic peptides and / or hormones has achieved limited success in the management of these diseases and conditions. In particular, dosage optimization is important for drugs and biologics used to treat metabolic diseases, especially those with narrow therapeutic windows. Hormones in general and peptides involved in glucose homeostasis often have narrow therapeutic windows. The narrow therapeutic window, coupled with the fact that such hormones and peptides typically have short half-lives (requiring frequent dosing to achieve clinical benefit), leads to difficulties in the management of these patients. Therefore, there remains a need for therapeutic agents with increased efficacy and safety in the treatment of metabolic diseases.

[0236] Therefore, one aspect of the present invention is to incorporate bioactive metabolic proteins relating to or used for the treatment of metabolic and cardiovascular diseases and conditions into a BPXTEN fusion protein to produce a composition effective in the treatment of such diseases and conditions. Metabolic proteins may include any protein having a biological, therapeutic, or preventative purpose or function that can be used to prevent, treat, mediate, or improve metabolic or cardiovascular diseases, conditions, or conditions. Table 4d provides a non-limiting list of such sequences of metabolic BPs covered by the BPXTEN fusion protein of the present invention. The metabolic protein of the BPXTEN composition of the present invention may be a protein that shows at least about 80% sequence identity with the protein sequences selected from Table 4d, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity.

[0237] Table 4d: Bioactive proteins used in metabolic disorders and cardiology

[0238]

[0239]

[0240]

[0241] "Anti-CD3" refers to monoclonal antibodies, species, sequence variants, and fragments thereof targeting the T-cell surface protein CD3, including OKT3 (also known as moromuzanab) and humanized anti-CD3 monoclonal antibodies (hOKT31(Ala-Ala)) (Herold et al., 2002, New England Journal of Medicine 346:1692-1698). The anti-CD3 fusion protein of the present invention can be particularly used to slow the onset of new-onset type 1 diabetes, including the use of anti-CD3 as a therapeutic effector and targeting moiety of a second therapeutic BP in BPXTEN compositions. The sequence of the variable region and the generation of anti-CD3 have been described in U.S. Patent Nos. 5,885,573 and 6,491,916.

[0242] "IL-1ra" refers to human IL-1 receptor antagonist proteins and their species and sequence variants possessing at least a portion of the biological activity of mature IL-1ra, including the sequence variant anaraxin. Anakinase is a non-glycosylated recombinant human IL-1ra, and differs from endogenous human IL-1ra by the addition of an N-terminal methionine residue. The commercial version of anakinase is... It is marketed. It binds to the IL-1 receptor with the same affinity as natural IL-1ra and IL-1b, but does not lead to receptor activation (signal transduction). The effect is attributed to the fact that IL-1ra has only one receptor-binding motif, compared to the two such motifs on IL-1α and IL-1β. Anahyperistin has 153 amino acids and a size of 17.3 kDa, and a reported half-life of approximately 4–6 hours.

[0243] Increased IL-1 production has been reported in patients suffering from various microbial infectious diseases and other conditions. The IL-1ra fusion protein of the present invention can be specifically used to treat any of the aforementioned diseases and conditions. IL-1ra has been cloned as described in U.S. Patent Nos. 5,075,222 and 6,858,409.

[0244] "Natriuretic peptide" refers to atrial natriuretic peptide (ANP), brain natriuretic peptide (BNP or B-type natriuretic peptide), and C-type natriuretic peptide (CNP); both human and non-human species and sequence variants possessing at least a portion of the biological activity of their mature paired natriuretic peptides. Sequences of useful forms of natriuretic peptides are disclosed in U.S. Patent Publication 20010027181. Examples of ANPs include human ANP (Kangawa et al., 1984, BBRC 118:131) or ANPs from various species, including porcine and rat ANPs (Kangawa et al., 1984, BBRC 121:585). Sequence analysis revealed that the BNP precursor originally consisted of 134 residues and was cleaved into a 108-amino acid precursor. A 32-amino acid cleavage from the C-terminus of the BNP precursor resulted in human BNP (77-108), which is the circulating physiologically active form. Human BNP, consisting of 32 amino acids, is involved in the formation of disulfide bonds (Sudoh et al., 1989, BBRC 159:1420) and is cited in US patents 5,114,923, 5,674,710, 5,674,710, and 5,948,761. BPXTENs containing one or more natriuretic functions can be used to treat hypertension, induce diuresis, induce natriuresis, diuresis or relaxation, vasodilation or relaxation, bind to natriuretic peptide receptors (e.g., NPR-A), inhibit aldosterone secretion from the adrenal glands, treat cardiovascular diseases and conditions, reduce, stop, or reverse cardiac remodeling after cardiac events or due to congestive heart failure, treat kidney diseases and conditions; treat or prevent ischemic stroke, and treat asthma.

[0245] "Heparin-binding growth factor 2" or "FGF-2" refers to the human FGF-2 protein and species and sequence variants of it that have at least a portion of the biological activity of their mature counterparts. FGF-2 has been cloned, as described in Burgess, WH and Maciag, T., Ann. Rev. Biochem., 58:575-606 (1989); Coulier, F. et al., 1994, Prog. Growth Factor Res. 5:1; and PCT Publication WO 87 / 01728.

[0246] "TNF receptor" refers to the human receptor for TNF and its biological receptor species and sequence variants that possess at least a portion of the activity of the mature TNFR. The X-ray crystal structure of the complex formed by the extracellular domain of the human p55 TNF receptor and TNFβ has been determined (Banner et al., 1993 Cell 73:431, which is incorporated herein by reference).

[0247] clotting factors

[0248] In hemophilia, blood clotting is disrupted due to the lack of certain plasma clotting factors. Human factor IX (FIX) is the proenzyme of a serine protease, which is an important component of the endogenous pathway of the coagulation cascade. Factor VIIa (FVIIa) protein has been found to be useful for treating bleeding episodes in hemophilia A or B patients with inhibitors against FVIII or FIX and in patients with acquired hemophilia, as well as for preventing bleeding during surgical interventions or invasive procedures in hemophilia A or B patients with inhibitors against FVIII or FIX. Therefore, there remains a need for a combination of factor IX and factor VIIa with an extended half-life and retained activity, when administered as part of a prophylactic and / or therapeutic regimen for hemophilia B, and for formulations with reduced side effects that can be administered via both intravenous and subcutaneous routes.

[0249] Coagulation factors included in the BPXTEN of this invention may include proteins having biological, therapeutic, or preventative purposes or functions that can be used to prevent, treat, mediate, or improve blood clotting disorders, diseases, or defects. Suitable coagulation proteins include bioactive polypeptides involved in the coagulation cascade as substrates, enzymes, or cofactors.

[0250] Table 4e provides a non-limiting list of coagulation factor sequences covered by the BPXTEN fusion protein of the present invention. Coagulation factors included in the BPXTEN of the present invention may be proteins that show at least about 80% sequence identity with protein sequences selected from Table 4e, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity.

[0251] Table 4e: Coagulation factor polypeptide sequences

[0252]

[0253]

[0254]

[0255]

[0256]

[0257]

[0258] "Factor IX" ("FIX") includes human factor IX protein and species and sequence variants of it having at least a portion of the biological receptor activity of mature factor IX. In some embodiments, a FIX peptide is a structural analog or peptide mimic of any FIX peptide described herein, including the sequence in Table 4e. In one specific example of the invention, FIX is human FIX. In another embodiment, FIX is a polypeptide sequence from Table 4e. Mature factor IX is a single-chain protein of 415 amino acid residues containing approximately 17% carbohydrates by weight (Schmidt 2003, Trends Cardiovasc Med, 13:39).

[0259] In some cases, a coagulation factor is factor IX, a sequence variant of factor IX, or a portion of factor IX, such as the exemplary sequence in Table 4e, as well as any protein or polypeptide substantially homologous to it, whose biological properties result in the activity of factor IX.

[0260] "Factor VII" (FVII) means human proteins and species and sequence variants of those having at least a portion of the biological activity of activating factor VII. Factor VII and recombinant human FVIIa have been introduced for uncontrolled bleeding in hemophiliac patients (with factor VIII or IX deficiency) who have developed inhibitors against alternative clotting factors. Recombinant human factor VIIa has utility in treating uncontrolled bleeding in hemophiliac patients (with factor VIII or IX deficiency), including those who have developed inhibitors against alternative clotting factors. In some embodiments, the FVII peptide is the activated form (FVIIa), a structural analog or peptide mimic of any FVII peptide described herein, including the sequences in Table 4e. Factor VII and VIIa have been cloned as described in U.S. Patent No. 6,806,063 and U.S. Patent Application No. 20080261886.

[0261] Growth hormone protein

[0262] "Growth hormone" or "GH" refers to the human growth hormone protein and its species and sequence variants, including but not limited to the 191-amino acid single-stranded human sequence of GH. This invention contemplates including in BPXTEN any GH homologous sequence, such as naturally occurring sequence fragments from primates, mammals (including domesticated animals), and non-natural sequence variants, that retain at least a portion of the biological activity or function of GH and / or can be used to prevent, treat, modulate, or improve GH-related diseases, defects, symptoms, or conditions. Non-mammalian GH sequences are well described in the literature. For example, sequence alignments of fish GH can be found in Genetics and Molecular Biology 2003, 26, pp. 295-300. Additionally, naturally occurring sequences homologous to human GH can be found using standard homology search techniques such as NCBI BLAST.

[0263] In one embodiment, the GH incorporated into a human or animal composition may be a recombinant polypeptide having a sequence corresponding to a protein found in nature. In another embodiment, the GH may be a sequence variant, fragment, homology, or mimic of a natural sequence that retains at least a portion of the biological activity of natural GH. Table 4f provides a non-limiting list of GH sequences from a wide variety of mammalian species covered by the BPXTEN fusion protein of the present invention. Any of these GH sequences or homologous derivatives constructed by rearranging individual mutations between species or families may be used in the fusion protein of the present invention. GH that can be incorporated into the BPXTEN fusion protein may include proteins that show at least about 80% sequence identity with the proteins selected from Table 4f, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity.

[0264] Table 4f: Amino acid sequences of growth hormones from animal species

[0265]

[0266]

[0267]

[0268]

[0269]

[0270]

[0271] Cytokines

[0272] BP can be a cytokine or one or more cytokines. Cytokines are proteins released by cells (such as chemokines, interferons, lymphokines, interleukins, and tumor necrosis factor) that can influence cellular behavior. Cytokines can be produced by a wide range of cells, including immune cells such as macrophages, B lymphocytes, T lymphocytes, and mast cells, as well as endothelial cells, fibroblasts, and various stromal cells. A given cytokine can be produced by more than one type of cell. Cytokines can be involved in producing systemic or local immune regulatory effects.

[0273] Some cytokines can act as pro-inflammatory cytokines. Pro-inflammatory cytokines are those involved in inducing or amplifying inflammatory responses. They can work with various cells of the immune system, such as neutrophils and leukocytes, to generate an immune response. Some cytokines can also act as anti-inflammatory cytokines. Anti-inflammatory cytokines are those involved in reducing inflammatory responses. In some cases, anti-inflammatory cytokines can modulate pro-inflammatory cytokine responses. Some cytokines can function as both pro-inflammatory and anti-inflammatory cytokines.

[0274] The cytokines encompassed by the compositions of this invention can be effective in treating a wide range of diseases, including but not limited to cancer, rheumatoid arthritis, multiple sclerosis, myasthenia gravis, systemic lupus erythematosus, Alzheimer's disease, schizophrenia, viral infections (e.g., chronic hepatitis C, AIDS), allergic asthma, retinal neurodegenerative processes, metabolic disorders, insulin resistance, and diabetic cardiomyopathy. Cytokines can be particularly useful in treating inflammatory and autoimmune conditions.

[0275] Examples of cytokines that can be regulated by the systems and compositions disclosed herein include, but are not limited to, lymphokines, monokines, and conventional polypeptide hormones other than human growth hormone. Cytokines include parathyroid hormone; thyroxine; insulin; proinsulin; relaxin; pro-relaxin; glycoprotein hormones such as follicle-stimulating hormone (FSH), thyroid-stimulating hormone (TSH), and luteinizing hormone (LH); liver growth factor; fibroblast growth factor; prolactin; placental prolactin; tumor necrosis factor-α; Müllerian duct inhibitory substance; mouse gonadotropin-related peptide; inhibin; activin; vascular endothelial growth factor; integrin; thrombopoietin (TPO); and neurotrophic factors. Growth factors, such as NGF-α; platelet-derived growth factor; transforming growth factor (TGF), such as TGF-α, TGF-β, TGF-β1, TGF-β2, and TGF-β3; insulin-like growth factor-I and-II; erythropoietin (EPO); Flt-3L; stem cell factor (SCF); bone-inducing factor; interferon (IFN), such as IFN-α, IFN-β, and IFN-γ; colony-stimulating factor (CSF), such as macrophage-CSF. (M-CSF); Granulocyte-Macrophage-CSF (GM-CSF); Granulocyte-CSF (G-CSF); Macrophage-Stimulating Factor (MSP); Interleukins (ILs), such as IL-1, IL-1a, IL-1b, IL-1RA, IL-18, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-9, IL-10, IL-11, IL-12, IL-12b, IL-13. IL-14, IL-15, IL-16, IL-17, and IL-20; tumor necrosis factors, such as CD154, LT-β, ​​TNF-α, TNF-β, 4-1BBL, APRIL, CD70, CD153, CD178, GITRL, LIGHT, OX40L, TALL-1, TRAIL, TWEAK, and TRANCE; and other polypeptide factors, including LIF, oncokinase M (OSM), and kit ligand (KL). Cytokine receptors are receptor proteins that bind cytokines. Cytokine receptors can be membrane-bound or soluble.

[0276] Target polynucleotides can encode cytokines. Non-restricted examples of cytokines include 4-1BBL, activin βA, activin βB, activin βC, activin βE, artemin (ARTN), BAFF / BLyS / TNFSF138, BMP10, BMP15, BMP2, BMP3, BMP4, BMP5, BMP6, BMP7, BMP8a, BMP8b, bone morphogenetic protein 1 (BMP1), CCL1 / TCA3, CCL11, CCL12 / MCP-5, CCL13 / MCP-4, CCL14, CCL15, CCL16, CCL17 / TARC, CCL18, CCL19, CCL2 / MCP-1, C... CL20, CCL21, CCL22 / MDC, CCL23, CCL24, CCL25, CCL26, CCL27, CCL28, CCL3, CCL3L3, CCL4, CCL4L1 / LAG-1, CCL5, CCL6, CCL7, CCL8, CCL9, CD153 / CD30L / TNFSF8, CD40L / CD154 / TNFSF5, CD40LG, CD70, CD70 / CD27L / TNFSF7, CLCF1, c-MPL / CD110 / TPOR, CNTF, CX3CL1, CXCL1, CXCL10, CXCL11, CXCL12, CXCL1 3. CXCL14, CXCL15, CXCL16, CXCL17, CXCL2 / MIP-2, CXCL3, CXCL4, CXCL5, CXCL6, CXCL7 / Ppbp, CXCL9, EDA-A1, FAM19A1, FAM19A2, FAM19A3, FAM19A4, FAM19A5, Fas ligand / FASLG / CD95L / CD178, GDF10, GDF11, GDF15, GDF2, GDF3, GDF4, GDF5, GDF6, GDF7, GDF8, GDF9, glial cell line-derived neurotrophic factor (GDNF), growth differentiation factor 1 (GDF1) IFNA1, IFNA10, IFNA13, IFNA14, IFNA2, IFNA4, IFNA5 / IFNaG, IFNA7, IFNA8, IFNB1, IFNE, IFNG, IFNZ, IFNω / IFNW1, IL11, IL18, IL18BP, IL1A, IL1B, IL1F10, IL1F3 / IL1RA, IL1F5, IL1F6, IL1F7, IL1F8, IL1F9, IL1RL2, IL31, IL33, IL6, IL8 / CXCL8, Inhibin-A, Inhibin-B, Leptin, LIF, LTA / TNFB / TNFSF1, LTB / TNFC,Neuronal rank protein (NRTN), OSM, OX-40L / TNFSF4 / CD252, persephin (PSPN), RANKL / OPGL / TNFSF11 (CD254), TL1A / TNFSF15, TNFA, TNF-α / TNFA, TNFSF10 / TRAIL / APO-2L (CD253), TNFSF12, TNFSF13, TNFSF14 / LIGHT / CD258, XCL1, and XCL2. In some embodiments, the target genes encode immune checkpoint inhibitors. Non-limiting examples of such immune checkpoint inhibitors include PD-1, CTLA-4, LAG3, TIM-3, A2AR, B7-H3, B7-H4, BTLA, IDO, KIR, and VISTA. In some embodiments, the target genes encode T cell receptor (TCR) α, β, γ, and / or δ chains.

[0277] In some cases, cytokines can be chemokines. Chemokines can be selected from, but are not limited to, the following group: ARMCX2, BCA-1 / CXCL13, CCL11, CCL12 / MCP-5, CCL13 / MCP-4, CCL15 / MIP-5 / MIP-1δ, CCL16 / HCC-4 / NCC4, CCL17 / TARC, CCL18 / PARC / MIP-4, CCL19 / MIP-3b, CCL2 / MCP-1, CCL20 / MIP-3α / MIP3A, CCL21 / 6Ckine, CCL22 / MDC, CCL23 / MIP 3. CCL24 / Eotaxin-2 / MPIF-2, CCL25 / TECK, CCL26 / Eotaxin-3, CCL27 / CTACK, CCL28, CCL3 / Mip1a, CCL4 / MIP1B, CCL4L1 / LAG-1, CCL5 / RANTES, CCL6 / C10, CCL8 / MCP-2, CCL9, CML5, CXCL1, CXCL10 / Crg-2, CXCL12 / SDF-1β, CXCL14 / BRAK, CXCL15 / Lungkine, CX CL16 / SR-PSOX, CXCL17, CXCL2 / MIP-2, CXCL3 / GROγ, CXCL4 / PF4, CXCL5, CXCL6 / GCP-2, CXCL9 / MIG, FAM19A1, FAM19A2, FAM19A3, FAM19A4 / TAFA4, FAM19A5, Fractalkine / CX3CL1, I-309 / CCL1 / TCA-3, IL-8 / CXCL8, MCP-3 / CCL7, NAP-2 / PPBP / CXCL7, XCL2 and Armo IL10.

[0278] Table 4g ​​provides a non-limiting list of such sequences of BP covered by the BPXTEN fusion protein of the present invention. The metabolic protein of the BPXTEN composition of the present invention can be a protein that shows at least about 80% sequence identity with the protein sequences selected from Table 4g, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity.

[0279] Table 4g: Cytokines used for conjugation

[0280]

[0281]

[0282] "IL-1ra" refers to human IL-1 receptor antagonist proteins and their species and sequence variants possessing at least a portion of the biological activity of mature IL-1ra, including the sequence variant anaraxin. Human IL-1ra is a mature glycoprotein of 152 amino acid residues. The IL-1ra-containing fusion protein of the present invention can be specifically used to treat any of the aforementioned diseases and conditions. IL-1ra has been cloned as described in U.S. Patent Nos. 5,075,222 and 6,858,409.

[0283] In some cases, BP can be IL-10. IL-10 can be a potent anti-inflammatory cytokine that inhibits the production of pro-inflammatory cytokines and chemokines. IL-10 can be used to treat autoimmune diseases and inflammatory conditions such as rheumatoid arthritis, multiple sclerosis, myasthenia gravis, systemic lupus erythematosus, Alzheimer's disease, schizophrenia, allergic asthma, retinal neurodegenerative processes, and diabetes.

[0284] In some cases, IL-10 can be modified to improve stability and reduce prolytic degradation. Modification can involve the substitution of one or more amide bonds. In some cases, one or more amide bonds within the IL-10 backbone can be substituted to achieve the aforementioned effects. One or more amide bonds (-CONH-) in IL-10 can be replaced with isosteric bonds that are amide bonds, such as -CH2NH-, -CH2S-, -CH2CH2-, -CH=CH- (cis and trans), -COCH2-, -CH(OH)CH2-, or -CH2SO-. Furthermore, the amide bonds in IL-10 can also be replaced by reduced isosteric pseudopeptide bonds. See Couder et al. (1993) Int. J. Peptide Protein Res. 41:181-184, which is incorporated herein by reference in its entirety.

[0285] One or more acidic amino acids, including aspartic acid, glutamic acid, homoglutamic acid, tyrosine, alkyl, aryl, arylalkyl and heteroarylsulfonamides of 2,4-diaminopropionic acid, ornithine or lysine and tetrazolium-substituted alkyl amino acids; and side-chain amide residues, such as asparagine, glutamine, and alkyl or aromatic substituted derivatives of asparagine or glutamine; and hydroxyl-containing amino acids, including serine, threonine, homoserine, 2,3-diaminopropionic acid, and alkyl or aromatic substituted derivatives of serine or threonine, which may be substituted.

[0286] One or more hydrophobic amino acids in IL-10, such as alanine, leucine, isoleucine, valine, ortholeucine, (S)-2-aminobutyric acid, (S)-cyclohexylalanine, or other simple α-amino acids, may be substituted with amino acids, including but not limited to aliphatic side chains from C1-C10 carbons, including branched, cyclic and linear alkyl, alkenyl, or alkynyl substitutions.

[0287] In some cases, one or more hydrophobic amino acids in IL-10 may be replaced, for example, by aromatic-substituted hydrophobic amino acids, including phenylalanine, tryptophan, tyrosine, sulfotyrosine, biphenylalanine, 1-naphthylalanine, 2-naphthylalanine, 2-benzothiophene alanine, 3-benzothiophene alanine, histidine, and amino, alkylamino, dialkylamino, aza, halogenated (fluorine, chlorine, bromine, or iodine), or alkoxy (C1-C4) substituted forms of the aromatic amino acids listed above. Illustrative examples include: 2-, 3- or 4-aminophenylalanine, 2-, 3- or 4-chlorophenylalanine, 2-, 3- or 4-methylphenylalanine, 2-, 3- or 4-methoxyphenylalanine, 5-amino-, 5-chloro-, 5-methyl- or 5-methoxytryptophan, 2'-, 3'- or 4'-amino-, 2'-, 3'- or 4'-chloro-, 2-, 3- or 4-biphenylalanine, 2'-, 3'- or 4'-methyl-, 2-, 3- or 4-biphenylalanine and 2- or 3-pyridylalanine;

[0288] One or more hydrophobic amino acids in IL-10, such as phenylalanine, tryptophan, tyrosine, sulfotyrosine, biphenylalanine, 1-naphthylalanine, 2-naphthylalanine, 2-benzothiophene alanine, 3-benzothiophene alanine, and histidine, including amino, alkylamino, dialkylamino, aza, halogenated (fluorine, chlorine, bromine, or iodine) or alkoxy groups, may be substituted with aromatic amino acids, including 2-, 3-, or 4-aminophenyl. Alanine, 2-, 3- or 4-chlorophenylalanine, 2-, 3- or 4-methylphenylalanine, 2-, 3- or 4-methoxyphenylalanine, 5-amino-, 5-chloro-, 5-methyl- or 5-methoxytryptophan, 2'-, 3'- or 4'-amino-, 2'-, 3'- or 4'-chloro-, 2-, 3- or 4-biphenylalanine, 2'-, 3'- or 4'-methyl-, 2-, 3- or 4-biphenylalanine and 2- or 3-pyridylalanine.

[0289] Amino acids containing basic side chains, including arginine, lysine, histidine, ornithine, 2,3-diaminopropionic acid, and homoarginine, may be substituted. Alkyl, alkenyl, or aryl substituted derivatives of the aforementioned amino acids may also be substituted. Examples include N-ε-isopropyl-lysine, 3-(4-tetrahydropyridyl)-glycine, 3-(4-tetrahydropyridyl)-alanine, N,N-γ,γ'-diethyl-homoarginine, α-methyl-arginine, α-methyl-2,3-diaminopropionic acid, α-methyl-histidine, and α-methyl-ornithine, wherein the alkyl group occupies the pre-R position of the α-carbon. Modified IL-10 may comprise an amide formed from any combination of the following: alkyl, aromatic, heteroaromatic, ornithine or 2,3-diaminopropionic acid, carboxylic acid, or any of many well-known activated derivatives, such as acyl chlorides, active esters, active azolides and related derivatives, lysine, and ornithine.

[0290] In some cases, IL-10 may contain one or more naturally occurring L-amino acids, synthetic L-amino acids, and / or D-enantiomers of amino acids. IL-10 polypeptides may contain one or more of the following amino acids: ω-aminodecanoic acid, ω-aminotetradecanoic acid, cyclohexylalanine, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, δ-aminovaleric acid, tert-butylalanine, tert-butylglycine, N-methylisoleucine, phenylglycine, cyclohexylalanine, leucine, naphthylalanine, ornithine, citrulline, 4-chlorophenylalanine, 2-fluorophenylalanine, pyridylalanine, 3-benzothioalanine, hydroxyproline, β-alanine, anthranilic acid, m-aminobenzoic acid, p-aminobenzoic acid, etc. The compounds include benzoic acid, m-aminomethylbenzoic acid, 2,3-diaminopropionic acid, α-aminoisobutyric acid, N-methylglycine (sarcosine), 3-fluorophenylalanine, 4-fluorophenylalanine, penicillamine, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, β-2-thiophene alanine, methionine sulfoxide, arginine, N-acetyllysine, 2,4-diaminobutyric acid, p-aminophenylalanine, N-methylvaline, homocysteine, homoserine, ε-aminohexanoic acid, ω-aminohexanoic acid, ω-aminoheptanoic acid, ω-aminooctanoic acid, and 2,3-diaminobutyric acid.

[0291] IL-10 may contain cysteine ​​residues or cysteine ​​groups, which may act as a linker to another peptide via disulfide bonding or provide for cyclization of the IL-10 polypeptide. Methods of introducing cysteine ​​or cysteine ​​analogues are known in the art; see, for example, U.S. Patent No. 8,067,532. The IL-10 polypeptide may be cyclized. Other cyclization methods include introducing an oxime linker or a lanthanine linker; see, for example, U.S. Patent No. 8,044,175. Any combination of amino acids (or non-amino acid moieties) that can form a cyclization bond may be used and / or introduced. The cyclization bond may be any combination of amino acids having a functional group (or amino acids and -(CH2)). n CO- or -(CH2) n The C6H4-CO- group is generated, and the functional group allows for the introduction of bridging. Some examples are disulfides, and disulfide analogs such as -(CH2). n -Carba bridges, thioacetals, thioether bridges (cystathionine or lanathionine), and bridges containing esters and ethers.

[0292] IL-10 can be substituted with N-alkyl, aryl, or main chain crosslinks to construct lactams and other cyclic structures, C-terminal hydroxymethyl derivatives, O-modified derivatives, and N-terminal modified derivatives, including substituted amides such as alkylamides and hydrazides. In some cases, IL-10 peptides are inverse analogs.

[0293] IL-10 can be a natural protein, a peptide fragment of IL-10 having at least a portion of the biological activity of natural IL-10, or a modified peptide. IL-10 can be modified to improve intracellular uptake. One such modification can be the attachment of a protein transduction domain. The protein transduction domain can be attached to the C-terminus of IL-10. Alternatively, the protein transduction domain can be attached to the N-terminus of IL-10. The protein transduction domain can be attached to IL-10 via a covalent bond. The protein transduction domain can be selected from any sequence listed in Table 4h.

[0294] Table 4h. Exemplary protein transduction domains

[0295] SEQ ID NO amino acid sequence 277 YGRKKRRQRRR; 278 RRQRRTSKLMKR 279 GWTLNSAGYLLGKINLKALAALAKKIL 280 KALAWEAKLAKALAKALAKHLAKALAKALKCEA 281 RQIKIWFQNRRMKWKK 282 YGRKKRRQRRR 283 RKKRRQRRR 284 YGRKKRRQRRR 285 RKKRRQRR 286 YARAAARQARA 287 THRLPRRRRRR 288 GGRRARRRRRR

[0296] The biosynthetic protein (BP) of human or animal compositions is not limited to natural full-length polypeptides, but also includes recombinant forms and biologically and / or pharmacologically active variants or fragments thereof. For example, those skilled in the art will understand that various amino acid substitutions can be prepared in the BP to produce variants without departing from the spirit of the invention regarding the biological activity or pharmacological properties of the BP. Examples of conserved substitutions of amino acids in the polypeptide sequence are shown in Table 5. However, in examples of BPXTENs in which the sequence identity of the BP is less than 100% compared to the specific sequence disclosed herein, the invention considers substitution for any of the other 19 natural L-amino acids for a given amino acid residue of a given BP, said amino acid residues may be located anywhere within the BP sequence, including adjacent amino acid residues. If any particular substitution results in an undesirable change in biological activity, the alternative amino acid can be employed, and the construct can be evaluated by the methods described herein, or by any techniques and guidelines regarding conserved and non-conserved mutations set forth, for example, in U.S. Patent No. 5,364,934, the contents of which are incorporated herein by reference in their entirety, or by methods generally known to those skilled in the art. Additionally, variants may include, for example, polypeptides in which one or more amino acid residues are added or deleted at the N-terminus or C-terminus of the full-length natural amino acid sequence of the BP, which retain at least a portion of the biological activity of the natural peptide.

[0297] Table 5: Exemplary Conserved Amino Acid Substitutions

[0298] Original residues Exemplary replacement Ala(A) val; leu; ile Arg(R) lys; gin; asn Asn(N) gin; his; Iys; arg Asp(D) glu Cys(C) ser Gln(Q) asn Glu(E) asp Gly(G) pro His(H) asn:gin:Iys:arg xIle(I) leu; val; met; ala; phe: leucine Leu(L) Leucine:ile:val;met;ala:phe Lys(K) arg:gin:asn Met(M) leu; phe; ile Phe(F) leu:val:ile;ala Pro(P) gly Ser(S) thr Thr(T) ser Trp(W) tyr Tyr(Y) trp:phe:thr:ser Val(V) ile; leu; met; phe; ala; leucine

[0299] In some embodiments, the BP incorporated into the BPXTEN polypeptide may have a sequence that shows at least about 80% sequence identity with the sequences from Tables 4a-4h, or alternatively at least about 81%, or about 82%, or about 83%, or about 84%, or about 85%, or about 86%, or about 87%, or about 88%, or about 89%, or about 90%, or about 91%, or about 92%, or about 93%, or about 94%, or about 95%, or about 96%, or about 97%, or about 98%, or about 99% or 100% sequence identity with the sequences from Tables 4a-4h. In some embodiments, the BP incorporated into BPXTEN may be a bispecific sequence comprising a first binding domain and a second binding domain, wherein the first binding domain has a specific binding affinity for tumor-specific markers or antigens of target cells, and the paired VL and VH sequences selected from anti-CD3 antibodies in Table 6f show at least about 80% sequence identity, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%... 96%, 97%, 98%, 99%, or 100% sequence identity; and wherein the second binding domain having specific binding affinity for effector cells shows at least about 80% sequence identity with the paired VL and VH sequences of the anti-target cell antibodies selected from Table 6a, or alternatively 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity. The activity of the BPs of the foregoing embodiments can be evaluated using the parameters as determined or measured or identified herein, and those sequences that retain at least about 40%, or about 50%, or about 55%, or about 60%, or about 70%, or about 80%, or about 90%, or about 95% or more of the activity compared to the corresponding native BP sequences are considered suitable for inclusion in human or animal BPXTENs. BPs found to retain suitable levels of activity can be linked to one or more XTEN peptides described above or anywhere else herein. In one embodiment, it was found that a BP retaining an appropriate level of activity could be linked to one or more XTEN peptides having at least about 80% sequence identity with sequences from Tables 3a-3b (e.g., at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity), resulting in a chimeric fusion protein.

[0300] T-cell binding agent

[0301] Other structural configurations of BPXTEN relate to XTEN-mediated protease-activated T-cell conjugates (“XPATs” or “XPATs”), wherein BP is a bispecific antibody (e.g., a bispecific T-cell conjugate). In some embodiments, the XPAT composition comprises a first portion including a first binding domain and a second binding domain, a second portion including a release segment, and a third portion including an XTEN-filling portion. In some embodiments, the XPAT composition has a configuration of formula Ia (described as N-terminus to C-terminus):

[0302] (Part 1) - (Part 2) - (Part 3) (Ia)

[0303] The first part comprises a bispecific scFv, wherein a first binding domain has a specific binding affinity for tumor-specific markers or antigens of target cells, and a second binding domain has a specific binding affinity for effector cells; the second part comprises a release segment (RS) that can be cleaved by a mammalian protease (which may be tumor-specific or antigen-specific, and thus activated) as described more fully below; and the third part is a filler portion. In the foregoing embodiments, the binding domains of the first part may be in the following order: (VL-VH)1-(VL-VH)2, where “1” and “2” represent the first binding domain and the second binding domain, respectively, or (VL-VH)1-(VH-VL)2, or (VH-VL)1-(VL-VH)2, or (VH-VL)1-(VH-VL)2, wherein the paired binding domains are linked by a polypeptide linker (as described more fully below). In one embodiment, substitutes for the first portions VL and VH are identified in Tables 6a-6f; substitutes for RS are identified in the sequences shown in Tables 8a-8b (described more fully below); and substitutes for the filler portion are identified herein by: XTEN; albumin-binding domain; albumin; IgG-binding domain; a polypeptide consisting of proline, serine, and alanine; fatty acid; Fc domain; polyethylene glycol (PEG), PLGA; and hydroxyethyl starch. When desired, the filler portion is XTEN, having at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the sequences identified by the sequences shown in Tables 3a-3b. In the foregoing embodiments, the composition is a recombinant fusion protein. In another embodiment, portions are linked by chemical conjugation.

[0304] In another embodiment, the XPAT composition has the configuration of formula IIa (described as N-terminus to C-terminus):

[0305] (Part Three) - (Part Two) - (Part One) (IIa)

[0306] The first part comprises a bispecific scFv, wherein a first binding domain has a specific binding affinity for tumor-specific markers or antigens of target cells, and a second binding domain has a specific binding affinity for effector cells; the second part comprises a release segment (RS) cleavable by mammalian proteases; and the third part is a filler portion. In the foregoing embodiments, the first binding domains may be in the order (VL-VH)1-(VL-VH)2, where “1” and “2” represent the first binding domain and the second binding domain, respectively, or (VL-VH)1-(VH-VL)2, or (VH-VL)1-(VL-VH)2, or (VH-VL)1-(VH-VL)2, wherein the paired binding domains are linked by a peptide linker described herein. In one embodiment, substitutes for the first portions VL and VH are identified in Tables 6a-6f; substitutes for RS are identified in the sequences shown in Tables 8a-8b; and substitutes for the filler portion are identified herein by: XTEN; albumin-binding domain; albumin; IgG-binding domain; a polypeptide consisting of proline, serine, and alanine; a fatty acid; and an Fc domain. When desired, the filler portion is XTEN, having at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a sequence selected from the sequences shown in Tables 3a-3b. In the foregoing embodiments, the composition is a recombinant fusion protein. In another embodiment, the portions are linked by chemical conjugation.

[0307] In another embodiment, the XPAT composition has a configuration of formula IIIa (described as N-terminus to C-terminus):

[0308] (Part 5) - (Part 4) - (Part 1) - (Part 2) - (Part 3)

[0309] (IIIa)

[0310] The first part is a bispecific scFv containing two scFvs, wherein the first binding domain has specific binding affinity for tumor-specific markers or antigens of target cells, and the second binding domain has specific binding affinity for effector cells; the second part contains a release segment (RS) cleavable by mammalian proteases; the third part is a filler portion; the fourth part contains a release segment (RS) cleavable by mammalian proteases, which may be the same as or different from the second part; and the fifth part is a filler portion, which may be the same as or different from the third part. In the foregoing embodiments, the binding domains of the first part may be in the order (VL-VH)1-(VL-VH)2, where “1” and “2” represent the first binding domain and the second binding domain, respectively, or (VL-VH)1-(VH-VL)2, or (VH-VL)1-(VL-VH)2, or (VH-VL)1-(VH-VL)2, wherein the paired binding domains are linked by a peptide linker described herein. In the foregoing embodiments, the substitutes for RS were identified in the sequences set forth in Tables 8a-8b. In the foregoing embodiments, the substitutes for the filler portion were identified herein by: XTEN; albumin-binding domain; albumin; IgG-binding domain; a polypeptide consisting of proline, serine, and alanine; a fatty acid; and an Fc domain. When desired, the filler portion is XTEN, having at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a sequence selected from the sequences shown in Tables 3a-3b. In the foregoing embodiments, the composition is a recombinant fusion protein. In another embodiment, the portion is linked by chemical conjugation.

[0311] Based on their design and specific components, human or animal compositions advantageously provide bispecific therapeutic agents that, once cleaved by proteases associated with target tissues or tissues unhealthy due to disease, exhibit higher selectivity, longer half-life, and result in less toxicity and fewer side effects. These human or animal compositions have an improved therapeutic index compared to bispecific antibody compositions known in the art. Such compositions can be used to treat certain diseases, including but not limited to cancers as described herein. Without being limited to any mechanistic theory, those skilled in the art will understand that the compositions of the invention achieve this reduction through a combination of mechanisms involving nonspecific interactions, including by positioning the binding domain to the steric hindrance of a large XTEN molecule, wherein the flexible, unstructured nature of the XTEN peptide, by tethering to the composition, enables it to oscillate and move around the binding domain, providing a blockage between the composition and tissue or cell, and by providing a reduced ability of the intact composition to penetrate cells or tissues due to its large molecular mass (contributed by both the actual molecular weight of the XTEN peptide and the large hydrodynamic radius of the unstructured XTEN peptide) compared to the size of the individual binding domains. However, in compositions designed in this way, when approaching a target tissue or cell carrying or secreting a protease capable of cleaving RS, or when the binding domain is internalized into the target cell or tissue after binding to a ligand, the bispecific binding domain is released from the bulk of the XTEN by the action of the protease, removing steric hindrance and allowing for greater freedom to exert its pharmacological effects. Human or animal compositions can be used to treat various conditions requiring selective delivery of the therapeutic bispecific antibody composition to cells, tissues, or organs. In one embodiment, the target tissue is cancer, which may be leukemia, lymphoma, or a tumor of an organ or system.

[0312] Combined structural domain

[0313] This disclosure considers the use of single-chain binding domains, such as, but not limited to, Fv, Fab, Fab', Fab'-SH, F(ab')2, linear antibodies, single-domain antibodies, single-domain camelid antibodies, single-chain antibody molecules (scFv), and biantibodies capable of binding to ligands or receptors associated with antigens of effector cells and diseased tissues or cells, said diseased tissues or cells being cancerous, tumors, or other malignant tissues. In some embodiments, a bispecific antibody comprises a first binding domain having binding specificity to a target cell marker and a second binding domain having binding specificity to an effector cell antigen. In some embodiments, the first and second binding domains may be non-antibody scaffolds, such as anticalin, adnectin, fynomer, affilin, affinity molecules, centyrins, DARPin. In other embodiments, the binding domain for tumor cell targets is a variable domain of a T-cell receptor modified to bind to the MHC of a peptide fragment loaded with a protein overexpressed by tumor cells. In some embodiments, XPAT compositions are designed with the following considerations in mind: the localization of the target cathepsin and the presence of the same protease in healthy tissue where it is not intended to be targeted, and the presence of the target ligand in healthy tissue, but a greater presence in unhealthy target tissue, in order to provide a wide therapeutic window. A “therapeutic window” refers to the maximum difference between the minimum effective dose and the maximum tolerated dose of a given therapeutic composition. To help achieve a wide therapeutic window, the binding domain of the first portion of the composition is shielded by the proximity of the filling portion (e.g., an XTEN peptide), wherein the binding affinity of the intact composition for one or both ligands is reduced compared to the composition cleaved by a mammalian protease, thereby releasing the first portion from the shielding effect of the filling portion.

[0314] Regarding the single-chain binding domain, as well established in the art, an Fv is the smallest antibody fragment containing a complete antigen recognition and binding site, consisting of a dimer of a heavy chain variable domain (VH) and a light chain variable domain (VL) bound non-covalently. Within each VH and VL chain are three complementarity-determining regions (CDRs) that interact to define the antigen-binding site on the surface of the VH-VL dimer; the six CDRs of the binding domain confer antigen-binding specificity to the antibody or the single-chain binding domain. In some cases, scFvs are generated where each binding domain has 3, 4, or 5 CDRs. The scaffold sequence flanking the CDRs has a tertiary structure that is substantially conserved across species in native immunoglobulins, and scaffold residues (FRs) act to hold the CDRs in their proper orientation. The constant domain is not required for binding function but helps stabilize the VH-VL interaction. In some embodiments, the domain of the polypeptide binding site can be a pair of VH-VL, VH-VH, or VL-VL domains from the same or different immunoglobulins; however, it is generally preferred to prepare single-chain binding domains using separate VH and VL chains from the parent antibody. The order of the VH and VL domains within the polypeptide chain is not limiting for the invention; the given order of domains can generally be reversed without loss of function, but it should be understood that the VH and VL domains are arranged in such a way that the antigen binding site can fold correctly. Therefore, the single-chain binding domains of bispecific scFv embodiments of human or animal compositions can be in the following order: (VL-VH) 1 -(VL-VH) 2 Where “1” and “2” represent the first and second binding structural domains, or (VL-VH), respectively. 1 -(VH-VL) 2 , or (VH-VL) 1 -(VL-VH) 2 , or (VH-VL) 1 -(VH-VL) 2 The paired binding domains are linked by peptide linkers as described below.

[0315] Therefore, the arrangement of the binding domains in the exemplary bispecific single-chain antibodies disclosed herein can be an arrangement in which the first binding domain is located at the C-terminus of the second binding domain. The arrangement of the V chains can be VH (target cell surface antigen)-VL (target cell surface antigen)-VL (effective cell antigen)-VH (effective cell antigen), VH (target cell surface antigen)-VL (target cell surface antigen)-VH (effective cell antigen)-VL (effective cell antigen), VL (target cell surface antigen)-VH (target cell surface antigen)-VL (effective cell antigen)-VH (effective cell antigen) or VL (target cell surface antigen)-VH (target cell surface antigen)-VL (effective cell antigen). For the arrangement where the second binding domain is located at the N-terminus of the first binding domain, the following orders are possible: VH(effective cell antigen)-VL(effective cell antigen)-VL(target cell surface antigen)-VH(target cell surface antigen), VH(effective cell antigen)-VL(effective cell antigen)-VH(target cell surface antigen)-VL(target cell surface antigen), VL(effective cell antigen)-VH(effective cell antigen)-VL(target cell surface antigen)-VH(target cell surface antigen) or VL(effective cell antigen)-VH(effective cell antigen)-VL(target cell surface antigen). As used herein, "its N-terminus" or "its C-terminus" and their grammatical variations indicate relative positioning within the primary amino acid sequence, rather than placement at the absolute N-terminus or C-terminus of the bispecific single-chain antibody. Therefore, as a non-limiting example, "the first binding domain located at the C-terminus of the second binding domain" means that the first binding is located on the carboxyl side of the second binding domain within the bispecific single-chain antibody, and does not preclude the possibility that other sequences, such as His-tags or other compounds, such as radioisotopes, are located at the C-terminus of the bispecific single-chain antibody.

[0316] In one embodiment, the chimeric peptide assembly composition includes a first portion comprising a first binding domain and a second binding domain, wherein each binding domain is an scFv, and each scFv comprises a VL and a VH. In another embodiment, the chimeric peptide assembly composition includes a first portion comprising a first binding domain and a second binding domain, wherein the binding domain is a biantibody configuration, and each domain comprises a VL domain and a VH. In the foregoing embodiments, the first domain has binding specificity for tumor-specific markers or antigens of target cells, and the second binding domain has binding specificity for effector cell antigens. In one of the foregoing embodiments, the effector cell antigen is expressed on or within effector cells. In one embodiment, the effector cell antigen is expressed on T cells such as CD4+, CD8+, or natural killer (NK) cells. In another embodiment, the effector cell antigen is expressed on B cells, mast cells, dendritic cells, or myeloid cells. In one embodiment, the effector cell antigen is CD3, the cluster 3 antigen of cytotoxic T cells. In some of the foregoing embodiments, the first binding domain exhibits binding specificity for tumor-specific markers associated with tumor cells. In one embodiment, the binding domain has binding affinity for tumor-specific markers, wherein the tumor cells may include, but are not limited to, cells derived from: stromal cell tumors, fibroblast tumors, myofibroblast tumors, glial cell tumors, epithelial cell tumors, adipocyte tumors, immune cell tumors, vascular cell tumors, and smooth muscle cell tumors.In one embodiment, the tumor-specific marker or target cell antigen may be α4 integrin, Ang2, B7-H3, B7-H6, CEACAM5, cMET, CTLA4, FOLR1, EpCAM, CCR5, CD19, HER2, HER2neu, HER3, HER4, HER1 (EGFR), PD-L1, PSMA, CEA, TROP-2, MUC1 (mucin), MUC-2, MUC3, MUC4, MUC5AC, MUC5B, MUC7, MUC16βhCG, Lewis-Y, CD20, CD33, CD38, CD30, CD56 (NCAM), CD133, ganglioside GD3; 9-O-acetyl-GD3, GM2, Globo H, fucose GM1, GD2, carbonic anhydrase IX, CD44v6, Nectin-4, or Sonic acid. Hedgehog (Shh), Wue-1, Plasma Cell Antigen 1, Chondroitin Sulfate Proteoglycan (MCSP), CCR8, Prostate 6-Transmembrane Epithelial Antigen (STEAP), Mesothelin, A33 Antigen, Prostate Stem Cell Antigen (PSCA), Ly-6, Desmosome Core Protein 4, Fetal Acetylcholine Receptor (fnAChR), CD25, Cancer Antigen 19-9 (CA19-9), Cancer Antigen 125 (CA-125), Müllerian Inhibitory Substance Receptor Type II (MISIIR), Sialized Tn Antigen (s The antibodies include TN, fibroblast activation antigen (FAP), endothelial sialic acid protein (CD248), epidermal growth factor receptor variant III (EGFRvIII), tumor-associated antigen L6 (TAL6), SAS, CD63, TAG72, Thomsen-Friedenreich antigen (TF antigen), insulin-like growth factor I receptor (IGF-IR), Cora antigen, CD7, CD22, CD70, CD79a, CD79b, G250, MT-MMP, F19 antigen, CA19-9, CA-125, alpha-fetoprotein (AFP), VEGFR1, VEGFR2, DLK1, SP17, ROR1, and EphA2. In one embodiment, the first binding domain exhibiting binding affinity for CD70 is its native ligand CD27, rather than the antibody fragment. In another embodiment, the first binding domain exhibiting binding affinity for B7-H6 is its native ligand Nkp30, rather than the antibody fragment.

[0317] The scFv embodiment of the XPAT composition of the present invention comprises a first binding domain and a second binding domain, wherein the VL and VH domains are derived from monoclonal antibodies that are specifically capable of binding to tumor-specific markers or target cell antigens and effector cell antigens, respectively. In other cases, the first binding domain and the second binding domain each comprise six CDRs derived from a monoclonal antibody that is specifically capable of binding to target cell markers such as tumor-specific markers and effector cell antigens. In other embodiments, the first binding domain and the second binding domain of the first portion of the human or animal composition may have 3, 4, or 5 CHRs within each binding domain. In other embodiments, embodiments of the present invention comprise a first binding domain and a second binding domain, wherein each comprises a CDR-H1 region, a CDR-H2 region, a CDR-H3 region, a CDR-L1 region, a CDR-L2 region, and a CDR-H3 region, wherein each of these regions is derived from a monoclonal antibody capable of binding to tumor-specific markers or target cell antigens and effector cell antigens. In one embodiment, the present invention provides a chimeric peptide assembly composition wherein the second binding domain comprises VH and VL regions derived from a monoclonal antibody capable of binding human CD3. In another embodiment, the present invention provides a chimeric peptide assembly composition wherein the scFv second binding domain comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity with or identical to the paired VL and VH sequences of the anti-CD3 antibodies shown in Table 6a. In another aspect, embodiments of the second domain of the present invention comprise CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, and CDR-H3 regions, wherein each of these regions is derived from a monoclonal antibody as shown in Table 6a. In the foregoing embodiments, the VH and / or VL domains can be configured as scFv, dual antibodies, single-domain antibodies, or single-domain camelid antibodies.

[0318] In other embodiments, the second domain of the human or animal composition is derived from an anti-CD3 antibody as shown in Table 6a. In one of the foregoing embodiments, the second domain of the human or animal composition comprises paired VL and VH region sequences of an anti-CD3 antibody as shown in Table 6a. In another embodiment, the present invention provides a chimeric polypeptide assembly composition wherein the second binding domain comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity with or identical to the paired VL and VH sequences of the huUCHT1 anti-CD3 antibody in Table 6a. In the foregoing embodiments, the VH and / or VL domains may be configured as scFv, part of a biantibody, a single-domain antibody, or a single-domain camelid antibody.

[0319] In other embodiments, the scFv of the first domain of the composition is derived from antitumor cell antibodies as shown in Table 6f. In another embodiment, the present invention provides a chimeric polypeptide assembly composition wherein the first binding domain comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity with or identical to the paired VL and VH sequences of the antitumor cell antibodies shown in Table 6f. In one of the foregoing embodiments, the first domain of the composition comprises the paired VL and VH region sequences of the antitumor cell antibodies disclosed herein. In the foregoing embodiments, the VH and / or VL domains may be configured as scFv, a portion of a biantibody, a single-domain antibody, or a single-domain camelid antibody.

[0320] In another embodiment, the chimeric polypeptide assembly composition includes a first portion comprising a first binding domain and a second binding domain, wherein the binding domain is a biantibody configuration, and each binding domain comprises a VL domain and a VH domain. In one embodiment, the biantibody embodiment of the present invention comprises a first binding domain and a second binding domain, wherein the VL and VH domains are derived from monoclonal antibodies having binding specificity to tumor-specific markers or target cell antigens and effector cell antigens, respectively. In another embodiment, the biantibody embodiment of the present invention comprises a first binding domain and a second binding domain, wherein each comprises a CDR-H1 region, a CDR-H2 region, a CDR-H3 region, a CDR-L1 region, a CDR-L2 region, and a CDR-H3 region, wherein each region is derived from a monoclonal antibody capable of binding to tumor-specific markers or target cell antigens and effector cell antigens. It is envisioned that the biantibody embodiment of the present invention comprises a first binding domain and a second binding domain, wherein the VL and VH domains are derived from monoclonal antibodies having binding specificity to tumor-specific markers or target cell antigens and effector cell antigens, respectively. In another aspect, the biantibody embodiments of the present invention comprise a first binding domain and a second binding domain, each comprising a CDR-H1 region, a CDR-H2 region, a CDR-H3 region, a CDR-L1 region, a CDR-L2 region, and a CDR-H3 region, wherein each of these regions is derived from a monoclonal antibody capable of binding tumor-specific markers or target cell antigens and effector cell antigens. In one embodiment, the present invention provides a chimeric polypeptide assembly composition wherein the second binding domain of the biantibody comprises paired VH and VL regions derived from a monoclonal antibody capable of binding human CD3. In another embodiment, the present invention provides a chimeric polypeptide assembly composition wherein the second binding domain of the biantibody comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity or similarity to the paired VL and VH sequences of anti-CD3 antibodies as shown in Table 6a. In another embodiment, the present invention provides a chimeric polypeptide assembly composition wherein the second binding domain of the biantibody comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity with or identical to the VL and VH sequences of the huUCHT1 antibody as shown in Table 6a. In other embodiments, the second binding domain of the biantibody in the composition is derived from the anti-CD3 antibody described herein.In another embodiment, the present invention provides a chimeric polypeptide assembly composition wherein the first binding domain of the biantibody comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity with or identical to the VL and VH sequences of the antitumor cell antibodies shown in Table 6f. In other embodiments, the first binding domain of the biantibody in the composition is derived from the antitumor cell antibodies described herein.

[0321] Therapeutic monoclonal antibodies from which VL and VH and CDR domains may be derived for use in human or animal compositions are known in the art. Sequences of the aforementioned antibodies can be obtained from publicly available databases, patents, or references. Furthermore, non-limiting examples of monoclonal antibodies and VH and VL sequences derived from anti-CD3 antibodies are described in Table 6a, and non-limiting examples of monoclonal antibodies targeting cancer, tumors, or target cell markers and VH and VL sequences are described in Table 6f.

[0322] Anti-CD3 binding domain

[0323] In some embodiments, the present invention provides a chimeric polypeptide assembly comprising a binding domain of a first portion having binding affinity for T cells. In one embodiment, the binding domain of a second portion comprises VL and VH sequences derived from a monoclonal antibody against a CD3 antigen. In another embodiment, the binding domain comprises VL and VH sequences derived from monoclonal antibodies against CD3ε and CD3δ. Monoclonal antibodies against CD3neu are known in the art. Exemplary, non-limiting examples of VL and VH sequences of monoclonal antibodies against CD3 are illustrated in Table 6a. In one embodiment, the present invention provides a chimeric polypeptide assembly comprising a binding domain having binding affinity for CD3, said binding domain comprising the anti-CD3 VL and VH sequences shown in Table 6a. In another embodiment, the present invention provides a chimeric polypeptide assembly comprising a binding domain of a first portion having binding affinity for CD3ε, said binding domain comprising the anti-CD3ε VL and VH sequences shown in Table 6a. In another embodiment, the present invention provides a chimeric peptide assembly composition wherein the second binding domain of the first portion of the scFv comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity with or identical to the paired VL and VH sequences of the huUCHT1 anti-CD3 antibody in Table 6a. In another embodiment, the present invention provides a chimeric peptide assembly composition comprising a binding domain having binding affinity for CD3, said binding domain comprising CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2, and CDR-H3 regions, wherein each is derived from the respective anti-CD3 VL and VH sequences shown in Table 6a. In another embodiment, the present invention provides a chimeric peptide assembly composition comprising a binding domain having binding affinity for CD3, the binding domain comprising CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2, and CDR-H3 regions, wherein the CDR sequences are RASQDIRNYLN (SEQ ID NO: 8034), YTSRLES (SEQ ID NO: 8035), QQGNTLPWT (SEQ ID NO: 8036), GYSFTGYTMN (SEQ ID NO: 8037), LINPYKGVST (SEQ ID NO: 8038), and SGYYGDSDWYFDV (SEQ ID NO: 8039).

[0324] The CD3 complex is a group of cell surface molecules that bind to the T-cell antigen receptor (TCR) and function in TCR cell surface expression and signal transduction cascades, which are initiated when a peptide:MHC ligand binds to the TCR. Typically, when an antigen binds to the T-cell receptor, CD3 signals through the cell membrane into the cytoplasm within the T cell. This leads to T-cell activation, which rapidly divides to produce sensitized new T cells to attack the specific antigens exposed to the TCR. The CD3 complex contains the CD3ε molecule, along with four other membrane-bound peptides (CD3-γ, -δ, -ζ, and -β). In humans, CD3-ε is encoded by the CD3E gene on chromosome 11. The intracellular domain of each CD3 chain contains an immune receptor tyrosine-based activation motif (ITAM), which acts as a nucleation site for intracellular signal transduction mechanisms following T-cell receptor binding.

[0325] Many therapeutic strategies modulate T-cell immunity by targeting TCR signaling, particularly anti-human CD3 monoclonal antibodies (mAbs) widely used in immunosuppressive regimens. The CD3-specific mouse mAb OKT3 was the first mAb approved for human use (Sgro, C. Side-effects of a monoclonal antibody, muromonab CD3 / orthoclone OKT3: bibliographic review. Toxicology 105:23-29, 1995), and is widely used clinically as an immunosuppressant in transplantation (Chatenoud, Clin. Transplant 7:422-430, (1993); Chatengoud, Nat. Rev. Immunol. 3:123-132 (2003); Kumar, Transplant. Proc. 30:1351-1352 (1998)), type 1 diabetes, and psoriasis. Importantly, anti-CD3 mAb can induce partial T cell signaling and clonal anergy (Smith, JA, Nonmitogenic Anti-CD3 Monoclonal Antibodies Deliver a Partial T Cell Receptor Signal and Induce Clonal Anergy J.Exp.Med.185:1413-1422(1997)). OKT3 has been described in the literature as a T cell mitogen and a potent T cell killer (Wong, JT. The mechanism of anti-CD3 monoclonal antibodies. Mediation of cytolysis by inter-T cell bridging. Transplantation 50:683-689(1990)). In particular, Wong's research confirmed that target killing can be achieved by bridging CD3 T cells and target cells, and that for bivalent anti-CD3MAB, neither FcR-mediated ADCC nor complement fixation is required to lyse target cells.

[0326] OKT3 exhibits time-dependent mitogenic and T-cell cytotoxic activity; following early activation of T cells leading to cytokine release, OKT3 subsequently blocks all known T-cell functions upon further administration. It is precisely because of this subsequent blockade of T-cell function that OKT3 has been found to have such widespread use as an immunosuppressant in therapeutic regimens aimed at reducing or even eliminating allogeneic tissue rejection. Other antibodies specific to the CD3 molecule are disclosed in Tunnacliffe, Int. Immunol. 1 (1989), 546-50. WO2005 / 118635 and WO2007 / 033230 describe anti-human monoclonal CD3ε antibodies. U.S. Patent 5,821,337 describes the VL and VH sequences of mouse anti-CD3 monoclonal Ab UCHT1 (muxCD3, Shalaby et al., J. Exp. Med. 175, 217-225 (1992) and a humanized variant of this antibody (hu UCHT1). U.S. Patent Application 20120034228 discloses a binding domain capable of binding to epitopes of human and non-chimpanzee primate CD3ε chains.

[0327] Table 6a: Anti-CD3 monoclonal antibodies and their sequences

[0328]

[0329]

[0330]

[0331]

[0332]

[0333]

[0334] *The underlined sequence (if it exists) is the CDR within VL and VH.

[0335] CD3 cell antigen binding fragment

[0336] In another aspect, this disclosure relates to antigen-binding fragments (AF2) having a specific binding affinity for effector cell antigens, which can be incorporated into any human or animal composition examples described herein. In some cases, the effector cell antigen is expressed on the surface of effector cells selected from plasma cells, T cells, B cells, cytokine-induced killer cells (CIK cells), mast cells, dendritic cells, regulatory T cells (RegT cells), helper T cells, myeloid cells, and NK cells.

[0337] Various AF2s that bind effector cell antigens have specific utility for pairing with antigen-binding fragments in compositional form, said antigen-binding fragments having binding affinity for EGFR antigens associated with diseased cells or tissues, in order to achieve cell killing of diseased cells or tissues. Binding specificity can be determined by complementarity-determining regions or CDRs, such as light chain CDRs or heavy chain CDRs. In many cases, binding specificity is determined by both light chain CDRs and heavy chain CDRs. A given combination of heavy chain CDRs and light chain CDRs provides a given binding pocket that imparts greater affinity and / or specificity to effector cell antigens compared to other reference antigens. A bispecific composition having a first antigen-binding fragment (AF1) against EGFR linked by a short, flexible peptide linker to a second antigen-binding fragment (AF2) that has binding specificity to effector cell antigens is bispecific, wherein each antigen-binding fragment has specific binding affinity to its respective ligand. Those skilled in the art will understand that in such compositions, AF1 against EGFR for diseased tissue is used in combination with AF2 against effector cell markers to bring effector cells close to cells of diseased tissue in order to achieve cell lysis of diseased tissue cells. Furthermore, AF1 and AF2 are incorporated into a specially designed polypeptide containing a cleavable release segment and XTEN to impart prodrug properties to the composition. When the composition is near diseased tissue having a protease capable of cleaving the release segment at one or more sites in the release segment sequence, the composition becomes activated by releasing the fused AF1 and AF2 after the release segment is cleaved.

[0338] In one embodiment, the AF2 of the human or animal composition has binding affinity for effector cell antigens expressed on the surface of T cells. In another embodiment, the AF2 of the human or animal composition has binding affinity for CD3. In yet another embodiment, the AF2 of the human or animal composition has binding affinity for members of the CD3 complex, said members comprising all known CD3 subunits of the CD3 complex, either individually or in independent combinations; for example, CD3ε, CD3δ, CD3γ, CD3ζ, CD3α, and CD3β. In yet another embodiment, AF2 has binding affinity for CD3ε, CD3δ, CD3γ, CD3ζ, CD3α, or CD3β.

[0339] The origin of the antigen-binding fragments considered in this disclosure may be derived from naturally occurring antibodies or fragments thereof, non-naturally occurring antibodies or fragments thereof, humanized antibodies or fragments thereof, synthetic antibodies or fragments thereof, hybrid antibodies or fragments thereof, or modified antibodies or fragments thereof. Methods for generating antibodies against a given target marker are well known in the art. For example, monoclonal antibodies may be prepared using a hybridoma method first described by Kohler et al., Nature, 256:495 (1975), or by a recombinant DNA method (US Patent No. 4,816,567). The structures of antibodies and their fragments, variable regions (VH and VL) of the antibody heavy and light chains, single-chain variable regions (scFv), complementarity-determining regions (CDRs), and domain antibodies (dAbs) are well understood. Methods for generating polypeptides having desired antigen-binding fragments with binding affinity for a given antigen are known in the art.

[0340] Those skilled in the art will understand that the use of the term "antigen-binding fragment" in the context of the compositions disclosed herein is intended to include a portion or fragment of an antibody that retains the ability to bind an antigen, which is a ligand of the corresponding intact antibody. In such embodiments, the antigen-binding fragment may be, but is not limited to, a CDR and intercalation framework region, a variable or hypervariable region of the antibody light chain and / or heavy chain (VL, VH), a variable fragment (Fv), a Fab' fragment, an F(ab')2 fragment, a Fab fragment, a single-chain antibody (scAb), a VHH camelid antibody, a single-chain variable fragment (scFv), a linear antibody, a single-domain antibody, a complementarity-determining region (CDR), a domain antibody (dAb), a single-domain heavy chain immunoglobulin of type BHH or BNAR, a single-domain light chain immunoglobulin, or other polypeptides known in the art containing an antibody fragment capable of binding an antigen. Antigen-binding fragments having CDR-H and CDR-L may be configured from the N-terminus to the C-terminus in a (CDR-H)-(CDR-L) or (CDR-H)-(CDR-L) orientation. The VL and VH of the two antigen-binding fragments can also be configured as single-chain biantibody configurations; that is, the VL and VH of AF1 and AF2 are configured with linkers of appropriate length to allow for biantibody arrangement.

[0341] The various CD3-binding AF2s of this disclosure have been specifically modified to enhance their stability in the peptide examples described herein. Antibody protein aggregation continues to be a significant issue for its developability and remains a major focus area in antibody production. Antibody aggregation can be triggered by the partial unfolding of its domains, leading to monomer-monomer binding, followed by nucleation and aggregate growth. Although the aggregation tendency of antibodies and antibody-based proteins can be influenced by external experimental conditions, they are strongly dependent on inherent antibody properties such as those determined by their sequence and structure. While it is well known that proteins are only slightly stable in their folded state, it is often less understood that most proteins are inherently prone to aggregation in their unfolded or partially unfolded state, and the resulting aggregates can be very stable and long-lived. Reduction in aggregation tendency has also been shown to be accompanied by increased expression titers, demonstrating that reducing protein aggregation is beneficial throughout the development process and can lead to more efficient clinical research pathways. For therapeutic proteins, aggregates are a significant risk factor for detrimental immune responses in patients and can form via a variety of mechanisms. Controlling aggregation can improve protein stability, manufacturability, loss rate, safety, formulation, titer, immunogenicity, and solubility. Intrinsic properties of proteins, such as size, hydrophobicity, electrostatics, and charge distribution, play a significant role in protein solubility. The low solubility of therapeutic proteins with hydrophobic surfaces has been shown to make formulation development more difficult and can lead to weak biodistribution in vivo, undesirable pharmacokinetic behavior, and immunogenicity. Reducing the overall surface hydrophobicity of candidate monoclonal antibodies can also provide benefits and cost savings related to purification and dosing regimens. Individual amino acids can be identified through structural analysis as having aggregation potential in antibodies and can be located in CDRs and framework regions. In particular, residues can be predicted to be at high risk of causing hydrophobicity problems in a given antibody. In one embodiment, this disclosure provides AF2 with the ability to specifically bind CD3, wherein the AF2 has at least one amino acid substitution of a hydrophobic amino acid in a framework region relative to a parent antibody or antibody fragment, wherein the hydrophobic amino acid is selected from isoleucine, leucine, or methionine. In another embodiment, CD3 AF2 has at least two amino acid substitutions of a hydrophobic amino acid in one or more framework regions, wherein the hydrophobic amino acid is selected from isoleucine, leucine, or methionine.

[0342] In designing the sequence of AF2 in the embodiments described herein, variations with respect to the net charge of the polypeptide were considered, particularly variations with respect to the antibody or antibody fragment constituting the specific embodiments of the invention set forth herein, wherein individual amino acid substitutions were prepared relative to the parent antibody used as a starting point. Related to these design considerations is the isoelectric point (pI) of the polypeptide, which is the pH at which the antibody or antibody fragment has no net charge. Antibodies or antibody fragments typically have a net positive charge, which tends to be associated with increased blood clearance and tissue retention, and generally a shorter half-life, while a net negative charge results in reduced tissue uptake and a longer half-life. This charge can be manipulated by mutations in the framework residues. The isoelectric point of the polypeptide can be determined arithmetically (e.g., computationally) or experimentally by in vitro determination. In some embodiments, the isoelectric points of AF1 and AF2 are designed within specific ranges of each other, thereby promoting stability.

[0343] In one embodiment, this disclosure provides AF2 for use in any of the peptide embodiments described herein, comprising CDR-L and CDR-H, wherein the AF2 (a) specifically binds to differentiation cluster 3 (CD3) of the T cell receptor; and (b) comprises CDR-H1, CDR-H2, and CDR-H3 having amino acid sequences having SEQ ID NO: 742, 743, and 744, respectively. In another embodiment, this disclosure provides AF2 for use in any of the peptide embodiments described herein, comprising CDR-L and CDR-H, wherein the AF2 (a) specifically binds to differentiation cluster 3 (CD3) of the T cell receptor; (b) comprises CDR-H1, CDR-H2, and CDR-H3 having amino acid sequences of SEQ ID NO: 742, 743, and 744, respectively; and (c) comprises CDR-L, wherein the CDR-L comprises CDR-L1 having an amino acid sequence of SEQ ID NO: 735 or 736, CDR-L2 having an amino acid sequence of SEQ ID NO: 738 or 739, and CDR-L3 having an amino acid sequence of SEQ ID NO: 740.In another embodiment, the aforementioned AF2 embodiment of this paragraph further comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises FR-L1 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO:746, FR-L2 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO:747, and ...0%, 91%, 92%, 93%, 9 The amino acid sequence of any one of NO:748-751 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity or is identical to that of FR-L3; the amino acid sequence of SEQ ID NO:754 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity or is identical to that of FR-L4; and the amino acid sequence of SEQ ID NO:755 or SEQ ID NO:754 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity or is identical to that of FR-L4. The amino acid sequence of NO:756 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-H1; the amino acid sequence of SEQ ID NO:759 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-H2; the amino acid sequence of SEQ ID NO:760 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-H3; and the amino acid sequence of SEQ ID NO:759 ... The amino acid sequence of NO:764 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or the same FR-H4.In another embodiment, AF2 used in any of the polypeptide embodiments described herein comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises FR-L1 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or equivalent to the amino acid sequence of SEQ ID NO:746, and FR-L2 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or equivalent to the amino acid sequence of SEQ ID NO:747 ...0%, 91%, 92%, 93 The amino acid sequence of NO:748 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent to FR-L3; the amino acid sequence of SEQ ID NO:754 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent to FR-L4; the amino acid sequence of SEQ ID NO:755 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent to FR-H1; and the amino acid sequence of SEQ ID NO:755 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent to FR-H1; and the amino acid sequence of SEQ ID NO:754 .... The amino acid sequence of NO:759 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent to FR-H2; the amino acid sequence of SEQ ID NO:760 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent to FR-H3; and the amino acid sequence of SEQ ID NO:764 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent to FR-H4.In another embodiment, AF2 used in any of the peptide embodiments described herein comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises FR-L1 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or equivalent to the amino acid sequence of SEQ ID NO:746, and FR-L2 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity or equivalent to the amino acid sequence of SEQ ID NO:747 ... The amino acid sequence of NO:749 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-L3; the amino acid sequence of SEQ ID NO:754 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-L4; the amino acid sequence of SEQ ID NO:755 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-H1; and the amino acid sequence of SEQ ID NO:755 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-H1; and the amino acid sequence of SEQ ID NO:754 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-H1. The amino acid sequence of ID NO:759 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H2; the amino acid sequence of SEQ ID NO:760 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H3; and the amino acid sequence of SEQ ID NO:764 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H4.In another embodiment, the AF2 of the human or animal polypeptide embodiment described herein comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises FR-L1 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO:746, FR-L2 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO:747, and ...0%, 91%, 92%, 93 The amino acid sequence of NO:750 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent FR-L3 with the amino acid sequence of SEQ ID NO:754; the amino acid sequence of SEQ ID NO:755 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent FR-L4 with the amino acid sequence of SEQ ID NO:755; and the amino acid sequence of SEQ ID NO:755 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or equivalent FR-H1 with the amino acid sequence of SEQ ID NO:754. The amino acid sequence of NO:759 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H2; the amino acid sequence of SEQ ID NO:760 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H3; and the amino acid sequence of SEQ ID NO:764 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H4.In another embodiment, the AF2 of the human or animal polypeptide embodiment described herein comprises a light chain framework region (FR-L) and a heavy chain framework region (FR-H), wherein AF2 comprises FR-L1 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO:746, FR-L2 showing at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO:747, and ...0%, 91%, 92%, 93 The amino acid sequence of NO:751 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-L3; the amino acid sequence of SEQ ID NO:754 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-L4; the amino acid sequence of SEQ ID NO:756 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is identical to FR-H1; and the amino acid sequence of SEQ ID NO:751 .... The amino acid sequence of NO:759 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H2; the amino acid sequence of SEQ ID NO:760 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H3; and the amino acid sequence of SEQ ID NO:764 shows at least 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity or is the same as FR-H4.

[0344] In another embodiment, this disclosure provides AF2 for use in any peptide embodiment described herein, wherein AF2 comprises a variable weight (VH) amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with or identical to the amino acid sequence of SEQ ID NO:766 or SEQ ID NO:769. In another embodiment, this disclosure provides AF2 for use in any peptide embodiment described herein, wherein AF2 comprises a variable light (VL) amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with or identical to the amino acid sequence of any one of SEQ ID NO:765, 767, 768, 770, or 771. In another embodiment, this disclosure provides AF2 for use in any of the polypeptide embodiments described herein, wherein the AF2 comprises a variable weight (VH) amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with or the same as the amino acid sequence of SEQ ID NO:766 or SEQ ID NO:769, and a variable light (VL) amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with or the same as the amino acid sequence of any one of SEQ ID NO:765, 767, 768, 770, or 771.

[0345] In another embodiment, this disclosure provides AF2 for use in any of the polypeptide embodiments described herein, wherein the AF2 comprises an amino acid sequence having at least 95%, 96%, 97%, 98%, or 99% sequence identity with or the same as an amino acid sequence of any of SEQ ID NO:776-780.

[0346] In another aspect, this disclosure provides AF2 antigen-binding fragments that bind to CD3 protein complexes, exhibiting enhanced stability compared to CD3-binding antibodies or antigen-binding fragments known in the art. Furthermore, the CD3 antigen-binding fragments of this disclosure are designed to confer a higher degree of stability to chimeric bispecific antigen-binding fragment compositions in which they are integrated, resulting in improved expression and recovery of the fusion protein, increased shelf life, and enhanced stability when administered to humans or animals. In one method, the CD3 AF2 of this disclosure is designed to have a higher degree of thermal stability compared to certain CD3-binding antibodies and antigen-binding fragments known in the art. As a result, CD3 AF2 used as a component of chimeric bispecific antigen-binding fragment compositions in which they are integrated exhibits advantageous pharmaceutical properties, including high thermal stability and low aggregation tendency, leading to improved expression and recovery during manufacturing and storage, and promoting a long serum half-life. Biophysical properties, such as thermal stability, are often limited by antibody variable domains, which vary considerably in their inherent properties. High thermal stability is often associated with high expression levels and other desirable properties, including less susceptibility to aggregation (Buchanan A et al., Engineering a therapeutic IgG molecule to address cysteinylation, aggregation and enhance thermal stability and expression. MAbs 2013; 5:255). Thermal stability is measured by the "melting temperature" (T0). m The melting temperature is determined by [the method of determination], defined as the temperature at which the next half of the molecules denatures. The melting temperature of each heterodimer is an indicator of its thermal stability. [The text then abruptly shifts to a different topic:] Determining T m The in vitro determination of the heterodimer is known in the art, including the methods described in the examples below. The melting point of the heterodimer can be measured using techniques such as differential scanning calorimetry (Chen et al. (2003) Pharm Res 20:1952-60; Ghirlando et al. (1999) Immunol Lett 68:47-52). Alternatively, the thermal stability of the heterodimer can be measured using circular dichroism (Murray et al. (2002) J. Chromatogr Sci 40:343-9) or as described in the examples below.

[0347] The thermal denaturation curves of the CD3-binding fragment and the anti-CD3 bispecific antibody comprising the said anti-CD3-binding fragment, combined with the reference of this disclosure, show that the constructs of this disclosure are more resistant to thermal denaturation than the antigen-binding fragment comprising the sequence shown in SEQ ID NO:781 or the control bispecific antibody (wherein the control bispecific antigen-binding fragment comprises SEQ ID NO:781) and the reference antigen-binding fragment in conjunction with the EGFR examples described herein. In one embodiment, the peptide of any human or animal composition example described herein comprises the anti-CD3 AF2 of the examples described herein, wherein the Tf of the antigen-binding fragment comprising the sequence of SEQ ID NO:781 is determined, as determined by the increase in melting temperature in an in vitro assay, to be more resistant to thermal denaturation. m In comparison, the T of AF2 m The temperature is at least 2°C higher, or at least 3°C ​​higher, or at least 4°C higher, or at least 5°C higher, or at least 6°C higher, or at least 7°C higher, or at least 8°C higher, or at least 9°C higher, or at least 10°C higher.

[0348] In another embodiment, the peptide of any human or animal composition embodiment described herein comprises AF2 that specifically binds to human or cynomolgus monkey CD3, with a dissociation constant (K... d The concentration is approximately 10 nM to approximately 400 nM, or approximately 50 nM to approximately 350 nM, or approximately 100 nM to approximately 300 nM, as determined in an in vitro antigen binding assay containing human or cynomolgus monkey CD3 antigen. In another embodiment, the polypeptide of any human or animal composition embodiment described herein comprises AF2 that specifically binds to human or cynomolgus monkey CD3, with a dissociation constant (K0). d Weaker than about 10 nM, or about 50 nM, or about 100 nM, or about 150 nM, or about 200 nM, or about 250 nM, or about 300 nM, or about 350 nM, or weaker than about 400 nM, as determined in an in vitro antigen binding assay. For clarity, K d The binding of the 400 antigen-binding fragment to its ligand is weaker than that of K. d The antigen-binding fragment is 10 nM. In another embodiment, the polypeptide of any human or animal composition embodiment described herein comprises AF2 that specifically binds to human or cynomolgus monkey CD3, with a binding affinity of up to 1 / 2, 1 / 3, 1 / 4, 1 / 5, 1 / 6, 1 / 7, 1 / 8, 1 / 9, or up to 1 / 10 of the antigen-binding fragment consisting of the amino acid sequence of SEQ ID NO:781, as determined by the respective dissociation constant (K0) in an in vitro antigen binding assay. dThe binding affinity for CD3 is determined by the specificity of the AF2-containing bispecific polypeptide, relative to the AF1 EGFR examples described herein incorporated into human or animal polypeptides, being at most 1 / 2, 1 / 3, 1 / 4, 1 / 5, 1 / 6, 1 / 7, 1 / 8, 1 / 9, 1 / 10, 1 / 20, 1 / 50, 1 / 100 or at most 1 / 1000, as determined by the respective dissociation constants (K0) in an in vitro antigen binding assay. d The binding affinity of a human or animal composition for a target ligand can be determined using the following methods: binding or competitive binding assays, such as the Biacore assay with a chip-based binding receptor or binding protein as described in U.S. Patent 5,534,617, or an ELISA assay, the assays described in the examples herein, a radioreceptor assay, or other assays known in the art. The binding affinity constant can then be determined using standard methods, such as the Scatchard analysis described by van Zoelen et al., Trends Pharmacol Sciences (1998) 19) 12:487, or other methods known in the art.

[0349] In a related aspect, this disclosure provides AF2, which binds to CD3 and is incorporated into a chimeric bispecific polypeptide composition designed to have an isoelectric point (pI) that imparts enhanced stability to the composition of this disclosure compared to corresponding compositions known in the art that contain CD3-binding antibodies or antigen-binding fragments. In one embodiment, the polypeptide of any human or animal composition embodiment described herein comprises CD3-binding AF2, wherein the AF2 exhibits a pI between 6.0 and 6.6, including the endpoints. In another embodiment, the polypeptide of any human or animal composition embodiment described herein comprises CD3-binding AF2, wherein the AF2 exhibits a pI at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0 pH units lower than the pI of a reference antigen-binding fragment comprising the sequence shown in SEQ ID NO:781. In another embodiment, the peptide of any human or animal composition embodiment described herein comprises an AF2 fused to an AF1 that binds to an EGFR antigen and binds to CD3, wherein the pI exhibited by the AF2 is within at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, or 1.5 pH units of the pI of the AF1 that binds to the EGFR antigen or its epitope. In another embodiment, the peptide of any human or animal composition embodiment described herein comprises an AF2 fused to an AF1 that binds to an EGFR antigen and binds to CD3, wherein the pI exhibited by the AF2 is within at least about 0.1 to about 1.5, or at least about 0.3 to about 1.2, or at least about 0.5 to about 1.0, or at least about 0.7 to about 0.9 pH units of the pI of the AF1. Specifically, this design, with the pIs of two of the antigen-binding fragments within such a range, is expected to confer a higher degree of stability to the chimeric bispecific antigen-binding fragment composition in which they are integrated, resulting in improved expression and enhanced recovery of the fusion protein in a soluble, non-aggregated form, increased shelf life of the formulated chimeric bispecific peptide composition, and enhanced stability when the composition is administered to humans or animals. In other words, placing AF2 and AF1 within a relatively narrow pI range allows for the selection of buffers or other solutions in which both AF2 and AF1 are stable, thereby promoting the overall stability of the composition.

[0350] In some embodiments, the VL and VH of the antigen-binding fragment are fused via a relatively long linker comprising 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 hydrophilic amino acids, which, when joined together, possess flexible properties. In one embodiment, the VL and VH of any scFv embodiment described herein are joined via a relatively long linker of hydrophilic amino acids, wherein the linker is GSGEGEGEGGEGEGEGEGSGEGEGEGEGGEGEGSG (SEQ ID NO: 8058).

[0351] TGSGEGSEGEGGGEGSEGEGSGEGGEGEGSGT (SEQ ID NO:8059),

[0352] GATPPETGAETESPGETTGGSAESEPPGEG (SEQ ID NO:8060, or

[0353] GSAAPTAGTTPSASPAPPTGGSSAAGSPST (SEQ ID NO: 8061). In another embodiment, AF1 and AF2 are linked together by a short linker of hydrophilic amino acids having 3, 4, 5, 6, or 7 amino acids. In one embodiment, the short linker sequence is SGGGGS (SEQ ID NO: 8062), GGGGS (SEQ ID NO: 8063), GGSGGS (SEQ ID NO: 8064), GGS, or GSP. In another embodiment, this disclosure provides a composition comprising a single-chain biantibody, wherein, upon folding, a first domain (VL or VH) pairs with a last domain (VH or VL) to form a scFv, and two intermediate domains pair to form another scFv, wherein the first and second domains, as well as the third and last domains, are fused together by one of the aforementioned short linkers, and the second and third variable domains are fused together by one of the aforementioned relatively long linkers. As those skilled in the art will understand, the selection of short and relatively long adapters is to prevent incorrect pairing of adjacent variable domains, thereby facilitating the formation of single-chain biantibody conformations of VL and VH containing a first antigen-binding fragment and a second antigen-binding fragment.

[0354] Table 6b. Exemplary CD3 CDR Sequences

[0355]

[0356] Table 6c. Exemplary CD3 FR sequences

[0357]

[0358]

[0359]

[0360] Table 6d: Exemplary VL and VH sequences

[0361]

[0362]

[0363] Table 6e: Exemplary scFv sequences

[0364]

[0365]

[0366]

[0367]

[0368]

[0369] Anti-EpCAM binding domain

[0370] In some embodiments, the present invention provides a chimeric peptide assembly composition comprising a binding domain having binding affinity for the tumor-specific marker EpCAM. In one embodiment, the binding domain comprises VL and VH sequences derived from a monoclonal antibody against EpCAM. Monoclonal antibodies against EpCAM are known in the art. Exemplary, non-limiting examples of EpCAM monoclonal antibodies and their VL and VH sequences are illustrated in Table 6f. In one embodiment, the present invention provides a chimeric peptide assembly composition comprising a binding domain having binding affinity for the tumor-specific marker EpCAM, said binding domain comprising the anti-EpCAM VL and VH sequences shown in Table 6f. In another embodiment, the present invention provides a chimeric peptide assembly composition wherein a first binding domain of the first portion comprises VH and VL regions, wherein each VH and VL region exhibits at least about 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identity or identicality to the paired VL and VH sequences of the 4D5MUCB anti-EpCAM antibody shown in Table 6f. In another embodiment, the present invention provides a chimeric peptide assembly composition comprising a binding domain having binding affinity for a tumor-specific marker, said binding domain comprising CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2, and CDR-H3 regions, wherein each is derived from the respective VL and VH sequences shown in Table 6f.

[0371] Table 6f. Anti-target cell monoclonal antibodies and their sequences

[0372]

[0373]

[0374]

[0375]

[0376]

[0377]

[0378]

[0379]

[0380]

[0381]

[0382]

[0383]

[0384]

[0385]

[0386]

[0387]

[0388]

[0389]

[0390]

[0391]

[0392]

[0393]

[0394]

[0395]

[0396]

[0397]

[0398]

[0399]

[0400]

[0401]

[0402]

[0403]

[0404]

[0405]

[0406]

[0407]

[0408]

[0409]

[0410]

[0411]

[0412]

[0413]

[0414]

[0415]

[0416]

[0417]

[0418]

[0419]

[0420]

[0421] *The underlined and bolded sequence (if it exists) is the CDR within VL and VH.

[0422] Epithelial cell adhesion molecule (EpCAM, also known as 17-1A antigen) is a 40-kDa membrane-integrated glycoprotein consisting of 314 amino acids expressed in some epithelial cells and many human cancers (see Balzar, The biology of the 17-1A antigen (Ep-CAM), J. Mol. Med. 1999, 77: 699-712). Due to their epithelial origin, tumor cells from most cancers express EpCAM on their surface (more than normal, healthy cells), including most primary, metastatic, and disseminated non-small cell lung cancer cells (Passlick, B. et al. The 17-1A antigen is expressed on primary, metastatic, and disseminated non-small cell lung carcinoma cells. Int. J. Cancer 87(4):548–552, 2000), gastric and gastroesophageal junction adenocarcinomas (Martin, IG. Expression of the 17-1A antigen in gastric and gastroesophageal junction adenocarcinomas: a potential immunotherapeutic target? J Clin Pathol 1999;52:701–704), and breast and colorectal cancers (Packeisen J et al. Detection of surface antigen 17-1A in breast and colorectal cancer. Hybridoma. 1999). 18(1):37-40), In breast cancer, overexpression of EpCAM on tumor cells is a predictor of survival (Gastl, Lancet. 2000, 356, 1981-1982). Due to their epithelial origin, tumor cells from most cancers express EpCAM on their surface.

[0423] In one embodiment, this document provides a bispecific chimeric polypeptide assembly composition having a first portion having an EpCAM-specific binding domain and a CD3-specific binding domain. The technical problem to be solved is to provide means and methods for generating improved compositions that exhibit good tolerability and more convenient (less frequent) administration properties for the effective treatment and / or improvement of oncological diseases. The solution to this technical problem is achieved through the embodiments disclosed herein and characterized in the claims.

[0424] Accordingly, in some embodiments, the present invention relates to chimeric peptide assembly compositions, wherein the compositions comprise a first portion comprising a bispecific single-chain antibody composition comprising at least two binding domains, wherein one of the domains binds an effector cell antigen such as a CD3 antigen, and a second domain binds an EpCAM antigen, wherein the binding domains comprise VL and VH specific to EpCAM and VL and VH specific to human CD3 antigen. Preferably, in embodiments, the binding domain specific to EpCAM has a value greater than 10. -7 Up to 10 -10 M of K d The value, as determined in an in vitro binding assay. In one of the foregoing embodiments, the binding domain is in the form of scFv. In another of the foregoing embodiments, the binding domain is in the form of a single-chain biantibody.

[0425] In some embodiments, the present invention provides a chimeric peptide assembly composition comprising a first binding domain having binding affinity for a tumor-specific marker and a second binding domain binding to an effector cell antigen, such as CD3 antigen. Tumor-specific markers constituting these embodiments of the present invention include, but are not limited to, CCR5, CD19, HER-2, HER-3, HER-4, EGFR, PSMA, CEA, MUC1, MUC2, MUC3, MUC4, MUC5AC, MUC5B, MUC7, βhCG, Lewis-Y, CD-20, CD33, CD30, ganglioside GD3, 9-O-acetyl-GD3, Globo H, fucose GM1, GD-2, carbonic anhydrase IX, CD44v6, and Sonic acid. Hedgehog, Wue-1, Plasma cell antigen 1, Chondroitin sulfate proteoglycan for melanoma, CCR8, Prostate 6-transmembrane epithelial antigen (STEAP), mesothelin, A33 antigen, Prostate stem cell antigen (PSCA), LY-6, SAS, Desmosome core protein 4, Fetal acetylcholine receptor, CD-25, Cancer antigen 19-9 (CA19-9), Cancer antigen 125 (CA-125), Müllerian inhibitory substance type II receptor (MISIIR), Sialidized Tn antigen, Fibroblast activation antigen (FAP), Endothelial sialic acid protein (CD248), Epidermal growth factor receptor variant III (EGFRvIII), Tumor-associated antigen L6 (TAL6), CD-63, TAG-72, Thomson-Frieden-Reich antigen (TF antigen), Insulin-like growth factor I receptor (IGF-IR), Cora antigen, CD7, CD22, CD79a, CD79b, G250, F19, EphA2, and MT-MM. In some embodiments, the present invention provides a chimeric peptide assembly composition comprising a first partial binding domain having binding affinity for a tumor-specific marker, the first partial binding domain comprising anti-marker VL and VH sequences. Exemplary, non-limiting examples of certain specific VL and VH sequences among these tumor markers are illustrated in Table 6f. In other embodiments, the present invention provides a chimeric peptide assembly composition comprising a first partial binding domain having binding affinity for a tumor-specific marker, the first partial binding domain comprising CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2, and CDR-H3 regions, each derived from a respective VL and VH sequence. Preferably, in embodiments, the binding has a binding affinity greater than 10. -7 Up to 10 -10 M of K d Values, such as those determined in in vitro binding assays.

[0426] In particular, it is considered that chimeric polypeptide assembly compositions may contain any of the aforementioned binding domains or sequence variants thereof, provided that the variants exhibit binding specificity for the antigen. In one embodiment, a sequence variant is generated by substituting amino acids in the VL or VH sequence with different amino acids. In deletion variants, one or more amino acid residues in the VL or VH sequence as described herein are removed. Thus, deletion variants comprise all fragments of the binding domain polypeptide sequence. In substitution variants, one or more amino acid residues of the VL or VH (or CDR) polypeptide are removed and replaced with substituted residues. In one aspect, substitution is conserved in nature, and this type of conserved substitution is well known in the art. Furthermore, it is particularly considered that compositions disclosed herein containing a first binding domain and a second binding domain can be used in any of the methods disclosed herein.

[0427] Unstructured conformation

[0428] Typically, despite the extended length of the polymer, the XTEN polypeptide components of the fusion proteins disclosed herein are designed to behave like denatured peptide sequences under physiological conditions. "Denatured" describes the state of the peptide in solution, characterized by a large degree of conformational freedom of the peptide backbone. Most peptides and proteins adopt a denatured conformation in the presence of high concentrations of denaturing agents or at high temperatures. Peptides in a denatured conformation exhibit, for example, a characteristic circular dichroism (CD) spectrum and are characterized by a lack of long-range interactions, as determined by NMR. "Denatured conformation" and "unstructured conformation" are used synonymously herein. In some cases, the present invention provides XTEN polypeptides that, under physiological conditions, can resemble denatured sequences that are largely lacking in secondary structure. In other cases, XTEN polypeptides may be substantially lacking in secondary structure under physiological conditions. As used in this context, "substantially lacking" means that less than 50% of the XTEN amino acid residues of each XTEN polypeptide contribute to the secondary structure, as measured or determined by the means described herein. As used in this context, “substantially lacking” means that at least about 60%, or about 70%, or about 80%, or about 90%, or about 95%, or at least about 99% of the XTEN amino acid residues in the XTEN sequence do not contribute to the secondary structure, as measured or determined by the means described herein.

[0429] Various methods have been established in the art to identify the presence or absence of secondary and tertiary structures in a given polypeptide. In particular, XTEN secondary structures can be measured spectrophotometrically, for example, by means of circular dichroism spectroscopy in the “far UV” spectral region (190–250 nm). Secondary structure elements, such as α-helices and β-sheets, each produce characteristic shapes and magnitudes of the CD spectrum. The secondary structure of polypeptide sequences can also be predicted via certain computer programs or algorithms, such as the well-known Chou-Fasman algorithm (Chou, PY et al. (1974) Biochemistry, 13:222–45) and the Garnier-Osguthorpe-Robson (“GOR”) algorithm (Garnier J, Gibrat JF, Robson B. (1996), GOR method for predicting protein secondary structure from amino acid sequence. Methods Enzymol 266:540–553), as described in U.S. Patent Application Publication No. 20030228309A1. For a given sequence, the algorithm can predict whether there is some secondary structure or no secondary structure at all, expressed as the total number and / or percentage of sequence residues that form, for example, α-helices or β-sheets, or predicted as the percentage of sequence residues that lead to the formation of random coils (which lack secondary structure).

[0430] In some cases, the XTEN peptide used in the fusion protein composition of the present invention may have an α-helix percentage ranging from 0% to less than about 5%, as determined by the Chou-Fasman algorithm. In other cases, the XTEN peptide constituting the fusion protein composition may have a β-sheet percentage ranging from 0% to less than about 5%, as determined by the Chou-Fasman algorithm. In some cases, the XTEN sequence of the fusion protein composition may have an α-helix percentage ranging from 0% to less than about 5% and a β-sheet percentage ranging from 0% to less than about 5%, as determined by the Chou-Fasman algorithm. In a preferred embodiment, the XTEN peptide constituting the fusion protein composition may have an α-helix percentage of less than about 2% and a β-sheet percentage of less than about 2%. In other cases, the XTEN sequence of the fusion protein composition may have a high random coil percentage, as determined by the GOR algorithm. In some embodiments, the XTEN peptide may have at least about 80%, more preferably at least about 90%, more preferably at least about 91%, more preferably at least about 92%, more preferably at least about 93%, more preferably at least about 94%, more preferably at least about 95%, more preferably at least about 96%, more preferably at least about 97%, more preferably at least about 98%, and most preferably at least about 99% random coils, as determined by the GOR algorithm.

[0431] Net charge

[0432] In other cases, XTEN peptides may possess unstructured properties conferred by incorporating amino acid residues with net charges and / or reducing the proportion of hydrophobic amino acids in the XTEN peptide. The overall net charge and net charge density can be controlled by modifying the content of charged amino acids in the XTEN peptide. In some cases, the net charge density of the XTEN in the composition may be greater than +0.1 or less than -0.1 charge / residue. In other cases, the net charge of the XTEN peptide may be about 0%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, or about 20% or more.

[0433] Because most tissues and surfaces in humans or animals have a net negative charge, XTEN peptides can be designed to have a net negative charge to minimize non-specific interactions between the XTEN peptide-containing composition and various surfaces such as blood vessels, healthy tissues, or various receptors. Unbound by any particular theory, XTEN peptides can adopt an open conformation due to electrostatic repulsion between the individual amino acids, which individually carry a high net negative charge and are distributed across the sequence of the XTEN peptide. Such a distribution of net negative charge in the extended sequence length of the XTEN peptide can lead to an unstructured conformation, which in turn can lead to an efficient increase in the hydrodynamic radius. Accordingly, in one embodiment, the present invention provides an XTEN peptide containing about 8, 10, 15, 20, 25, or even about 30% glutamic acid. The XTEN peptides of the compositions of the present invention generally have little or no positively charged amino acids. In some cases, the XTEN peptide may have less than about 10% positively charged amino acid residues, or less than about 7%, or less than about 5%, or less than about 2% positively charged amino acid residues. However, the present invention contemplates constructs in which a limited number of positively charged amino acids, such as lysine, can be incorporated into the XTEN polypeptide to allow conjugation between the ε-amine of lysine and reactive groups on the peptide, linker bridges, or reactive groups on a drug or small molecule to be conjugated to the XTEN polypeptide backbone. As described above, fusion proteins can be constructed comprising one or more XTEN polypeptides, a bioactive protein, and a chemotherapeutic agent for treating metabolic diseases or conditions, wherein the maximum number of molecules of the agent incorporated into the XTEN polypeptide component is determined by the number of lysine or other amino acids (e.g., cysteine) with reactive side chains incorporated into the XTEN.

[0434] In some cases, XTEN peptides may contain charged residues separated by other residues such as serine or glycine, which can lead to better expression or purification behavior. Based on net charge, XTEN peptides in human or animal compositions may have isoelectric points (pI) of 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, or even 6.5. In preferred embodiments, the XTEN peptide has an isoelectric point of 1.5 to 4.5. In these embodiments, XTEN incorporated into the BPXTEN fusion protein compositions of the present invention will carry a net negative charge under physiological conditions, which can promote the unstructured conformation of the XTEN peptide component and reduced binding to mammalian proteins and tissues.

[0435] Since hydrophobic amino acids can endow peptides with structure, the present invention provides that the content of hydrophobic amino acids in XTEN peptides is typically less than 5%, or less than 2%, or less than 1%. In one embodiment, the amino acid content of methionine and tryptophan in the XTEN component of the BPXTEN fusion protein is typically less than 5%, or less than 2%, and most preferably less than 1%. In another embodiment, the XTEN peptide has a sequence having less than 10% positively charged amino acid residues, or less than about 7%, or less than about 5%, or less than about 2% positively charged amino acid residues, the sum of methionine and tryptophan residues being less than 2%, and the sum of asparagine and glutamine residues being less than 10% of the total XTEN peptide.

[0436] Increased hydrodynamic radius

[0437] In some embodiments, the XTEN peptide may have a high hydrodynamic radius, conferring a correspondingly increased apparent molecular weight to the BPXTEN fusion protein incorporating the XTEN peptide. The linking of the XTEN peptide to the BP sequence can result in a BPXTEN composition having an increased hydrodynamic radius, an increased apparent molecular weight, and an increased apparent molecular weight factor compared to BP not linked to the XTEN peptide. For example, in therapeutic applications where a prolonged half-life is required, incorporating an XTEN peptide with a high hydrodynamic radius into a composition comprising one or more BP fusion proteins can effectively expand the hydrodynamic radius of the composition beyond approximately 3-5 nm of glomerular pore size (corresponding to an apparent molecular weight of approximately 70 kDa) (Caliceti. 2003. Pharmacokinetic and biodistribution properties of poly(ethylene glycol)-protein conjugates. Adv. Drug Deliv. Rev. 55: 1261-1277), resulting in a reduced renal clearance of circulating proteins. Unbound by any particular theory, XTEN peptides can adopt open conformations due to electrostatic repulsion between the individual charges of the peptide or the inherent flexibility conferred by specific amino acids in the sequence that lack the potential to confer secondary structure. The open, extended, and unstructured conformations of XTEN peptides can have a larger proportionate hydrodynamic radius compared to peptides with secondary and / or tertiary structures, such as typical globular proteins, having comparable sequence lengths and / or molecular weights. Methods for determining the hydrodynamic radius are well known in the art, for example, by using size exclusion chromatography (SEC), as described in U.S. Patent Nos. 6,406,632 and 7,294,513. The addition of an increased length of XTEN peptide results in a proportional increase in the parameters of hydrodynamic radius, apparent molecular weight, and apparent molecular weight factor, allowing BPXTEN to be customized for desired characteristic truncation of apparent molecular weight or hydrodynamic radius. Accordingly, in some embodiments, BPXTEN fusion proteins can be configured with XTEN peptides, wherein the fusion protein can have a hydrodynamic radius of at least about 5 nm, or at least about 8 nm, or at least about 10 nm, or 12 nm, or at least about 15 nm. In the foregoing embodiments, the large hydrodynamic radius imparted by the XTEN peptide in the BPXTEN fusion protein can lead to a decrease in the renal clearance of the resulting fusion protein, resulting in a corresponding increase in the terminal half-life, an increase in the mean residence time, and / or a decrease in renal clearance.

[0438] In another embodiment, XTEN peptides of selected lengths and sequences may be selectively incorporated into BPXTEN to produce a fusion protein having the following apparent molecular weights under physiological conditions: at least about 150 kDa, or at least about 300 kDa, or at least about 400 kDa, or at least about 500 kDa, or at least about 600 kDa, or at least about 700 kDa, or at least about 800 kDa, or at least about 900 kDa, or at least about 1000 kDa, or at least about 1200 kDa, or at least about 1500 kDa, or at least about 1800 kDa, or at least about 2000 kDa, or at least about 2300 kDa or more. In another embodiment, an XTEN peptide of selected length and sequence may be selectively linked to BP to result in a BPXTEN fusion protein having an apparent molecular weight factor of at least three, alternatively at least four, alternatively at least five, alternatively at least six, alternatively at least eight, alternatively at least 10, alternatively at least 15, or at least 20 or greater under physiological conditions. In another embodiment, the BPXTEN fusion protein has an apparent molecular weight factor of about 4 to about 20, or about 6 to about 15, or about 8 to about 12, or about 9 to about 10 relative to the actual molecular weight of the fusion protein under physiological conditions. In some embodiments, the (fusion) peptide exhibits an apparent molecular weight factor greater than about 6 under physiological conditions.

[0439] Increased terminal half-life

[0440] In some embodiments, the (fusion) peptide has at least two, three, four, or five times the terminal half-life compared to a bioactive peptide not linked to an XTEN peptide.

[0441] Administering any embodiment of the BPXTEN fusion protein described herein at a therapeutically effective dose to a human or animal in need may result in at least two, three, four, five, or more times the time spent within the therapeutic window of the fusion protein compared to a corresponding BP not linked to the XTEN peptide and administered to a human or animal at a comparable dose.

[0442] low immunogenicity

[0443] In another aspect, the present invention provides compositions wherein the XTEN polypeptide has low immunogenicity or is substantially non-immunogenic. Several factors can contribute to the low immunogenicity of the XTEN polypeptide, such as a substantially non-repetitive sequence, its unstructured conformation, high solubility, low degree or absence of self-aggregation, low degree or absence of protease-cleaving sites within the sequence, and low degree or absence of epitopes in the XTEN polypeptide.

[0444] Those skilled in the art will understand that, in general, polypeptides having highly repetitive short amino acid sequences (e.g., sequences of 200 amino acids or more containing an average of 20 or more repeats of a limited set of 3- or 4-mers) and / or having continuously repeating amino acid residues (e.g., 5- or 6-mer sequences having the same amino acid residues) tend to aggregate or form higher-order structures or form contacts, resulting in crystalline or pseudocrystalline structures.

[0445] In some embodiments, the XTEN polypeptide is substantially non-repetitive, wherein the XTEN amino acid sequence does not have three consecutive amino acids of the same amino acid type, unless the amino acid is serine, in which case no more than three consecutive amino acids may be serine residues; and wherein the XTEN amino acid sequence does not contain a 3-amino acid sequence (3-mer) appearing more than 16, 14, 12, or 10 times within a 200-amino acid sequence of the XTEN polypeptide. Those skilled in the art will understand that such substantially non-repetitive sequences have a lower tendency to aggregate, and therefore, it is possible to design long XTEN sequences with a relatively low frequency of charged amino acids, which are more likely to aggregate if the sequence or amino acid residues are otherwise more repetitive.

[0446] Conformational epitopes are regions formed on the surface of proteins, consisting of multiple discontinuous amino acid sequences of protein antigens. Precise protein folding allows these sequences to achieve well-defined, stable spatial conformations, or epitopes, which can be recognized as "exogenous" by the host's humoral immune system, leading to the production of antibodies against the protein or triggering a cell-mediated immune response. In the latter case, the immune response against the protein in an individual is heavily influenced by T cell epitope recognition, a function specific to the binding of an individual's HLA-DR allotype peptide. MHC class II peptide complexes can induce activation within T cells through binding to homologous T cell receptors on the surface of T cells, along with cross-binding to certain other co-receptors such as CD4 molecules. Activation leads to the release of cytokines, further activating other lymphocytes, such as B cells, to produce antibodies, or activating T killer cells as a full cellular immune response.

[0447] The ability of a peptide to bind a given MHC class II molecule for presentation on the surface of an APC (antigen-presenting cell) depends on many factors; most notably its primary sequence. In one embodiment, a lower degree of immunogenicity can be achieved by designing an XTEN peptide resistant to antigen processing in antigen-presenting cells and / or selecting sequences that do not bind well to MHC receptors. The present invention provides a BPXTEN fusion protein having a substantially non-repetitive XTEN peptide designed to reduce binding to MHC II receptors and avoid the formation of epitopes related to T-cell receptor or antibody binding, resulting in a lower degree of immunogenicity. This avoidance of immunogenicity is partly a direct result of the conformational flexibility of the XTEN peptide; i.e., a lack of secondary structure due to the selection and order of amino acid residues. For example, sequences with a low tendency to adapt to a tight-folded conformation are of particular interest in aqueous solutions or under physiological conditions that can lead to conformational epitopes. Administration of fusion proteins containing XTEN peptides generally does not lead to the formation of neutralizing antibodies against the XTEN peptide using conventional therapeutic practices and dosing, and can also reduce the immunogenicity of BP fusion couplers in BPXTEN compositions.

[0448] In one embodiment, the XTEN peptide for use in human or animal fusion proteins may be substantially free of epitopes recognized by human T cells. Elimination of such epitopes for the purpose of generating proteins with lower immunogenicity has been previously disclosed; see, for example, WO 98 / 52976, WO 02 / 079232 and WO 00 / 3317, which are incorporated herein by reference. Assays for human T cell epitopes have been described (Stickler, M. et al. (2003) J Immunol Methods, 281:95-108). Of particular interest are peptide sequences that can be oligomerized without generating T cell epitopes or non-human sequences. This can be achieved by testing for direct repetition of the sequences in relation to the presence of T cell epitopes and the appearance of non-human 6- to 15-mer, and especially 9-mer, sequences, and then modifying the design of the XTEN peptide to eliminate or disrupt the epitope sequence. In some cases, the XTEN peptide is substantially non-immunogenic by limiting the number of epitopes predicted to bind to the MHC receptor. With a decrease in the number of epitopes capable of binding MHC receptors, there is a potential for a decline in T cell activation and a corresponding decrease in T cell helper function, reduced B cell activation or upregulation, and reduced antibody production. The low degree of predicted T cell epitopes can be determined by epitope prediction algorithms, such as TEPITOPE (Sturniolo, T. et al. (1999) Nat Biotechnol, 17:555-61), as illustrated in Example 74 of International Patent Application Publication No. WO 2010 / 144502 A2, which is incorporated herein by reference in its entirety. As disclosed in Sturniolo, T. et al. (1999) Nature Biotechnology 17:555), the TEPITOPE score for a given peptide framework within a protein is the K-value of that peptide framework binding to multiple of the most common human MHC alleles. d The logarithm of (dissociation constant, affinity, dissociation rate). The scoring range exceeds at least 20 logarithms, approximately 10 to approximately -10 (corresponding to 10e). 10 K d up to 10e -10 K d (binding constraints), and can be reduced by avoiding hydrophobic amino acids that can act as anchoring residues during peptide display on the MHC, such as M, I, L, V, and F. In some embodiments, the XTEN peptide incorporated into BPXTEN does not have predictive T cell epitopes with a TEPIE score of about -5 or greater, or -6 or greater, or -7 or greater, or -8 or greater, or a TEPIE score of -9 or greater. As used herein, a score of "-9 or greater" will encompass TEPIE scores from 10 to -9, including the endpoints, but will not encompass scores of -10, as -10 is less than -9.

[0449] In another embodiment, the XTEN peptides of the present invention, including those incorporated into human or animal BPXTEN fusion proteins, can be substantially non-immunogenic by restricting known proteolytic sites from the XTEN peptide sequence, reducing the processing of the XTEN peptide into small peptides that can bind to MHC II receptors. In another embodiment, the XTEN peptides can be substantially non-immunogenic by using sequences that are substantially lacking in secondary structure, conferring resistance to many proteases due to the high entropy of the structure. Accordingly, a reduced TEPITPE score and the elimination of known proteolytic sites from the XTEN peptides can cause XTEN-peptide compositions, including XTEN peptides of BPXTEN fusion protein compositions, to be substantially non-binding to mammalian receptors, including receptors of the immune system. In one embodiment, the XTEN peptide of the BPXTEN fusion protein may have a >100 nM K binding affinity to mammalian cell surface or circulating peptide receptors. d or greater than 500 nM K d or greater than 1 μM K d The combination of.

[0450] Furthermore, the substantially non-repetitive sequences and corresponding lack of epitopes in such embodiments of the XTEN peptide can limit the ability of B cells to bind to or be activated by the XTEN peptide. While the XTEN peptide can contact many different B cells along its extended sequence, each individual B cell can only have one or a small number of contacts with the individual XTEN peptide. As a result, the XTEN peptide can generally have a much lower tendency to stimulate B cell proliferation and thus stimulate an immune response. In one embodiment, BPXTEN can have reduced immunogenicity compared to the corresponding unfused BP. In one embodiment, administration of up to three parenteral doses of BPXTEN to a mammal can result in detectable anti-BPXTEN IgG at a serum dilution of 1:100, rather than at a dilution of 1:1000. In another embodiment, administration of up to three parenteral doses of BPXTEN to a mammal can result in detectable anti-BP IgG at a serum dilution of 1:100, rather than at a dilution of 1:1000. In another embodiment, administration of up to three parenteral doses of BPXTEN to a mammal can result in detectable anti-XTEN IgG at a serum dilution of 1:100, rather than at a dilution of 1:1000. In the foregoing embodiments, the mammal may be a mouse, rat, rabbit, or cynomolgus monkey.

[0451] Compared to sequences with fewer non-repetitive sequences (e.g., sequences with the same three consecutive amino acids), an additional characteristic of certain embodiments of XTEN peptides with substantially non-repetitive sequences may be that the non-repetitive XTEN peptides form a weaker contact (e.g., monovalent interaction) with the antibody, resulting in a lower likelihood of immune clearance, wherein the BPXTEN composition can be retained for an increased period of time in the cycle.

[0452] In some embodiments, the (fusion) peptide is less immunogenic than a bioactive peptide not linked to the XTEN peptide, wherein immunogenicity is determined by measuring the production of IgG antibodies that selectively bind to the bioactive peptide after administration of a comparable dose to a human or animal.

[0453] Interval zone and BP release section

[0454] In some embodiments, at least a portion of the bioactivity of each BP is retained by the intact BPXTEN. In some embodiments, the BP component becomes bioactive or has increased bioactivity after being released from the XTEN polypeptide by cleavage of an optional cleavage sequence incorporated into the spacer region sequence within the BPXTEN, as described more fully below.

[0455] Any set of spacer region sequences is optional in fusion proteins covered by this invention. Spacer regions can be provided to enhance the expression of fusion proteins derived from host cells or to reduce steric hindrance, wherein the BP component can take its desired tertiary structure and / or appropriately interact with its target molecules. For spacer regions and methods for identifying desired spacer regions, see, for example, George et al. (2003) Protein Engineering 15:871-879, which is specifically incorporated herein by reference. In one embodiment, the spacer region comprises one or more peptide sequences of 1 to 50 amino acid residues, or about 1 to 25 residues, or about 1 to 10 residues. Excluding cleavage sites, the spacer region sequence may comprise any of 20 native L amino acids and preferably comprises sterically unhindered hydrophilic amino acids, which may include, but are not limited to, glycine (G), alanine (A), serine (S), threonine (T), glutamate (E), and proline (P). In some embodiments, the spacer region may be polyglycine or polyalanine, or a mixture of predominantly glycine and alanine residues. The spacer region polypeptides excluding the cleavage sequence largely lack secondary structure. In one embodiment, one or both spacer region sequences in the BPXTEN fusion protein composition may also each contain a cleavage sequence, which may be the same or different, said cleavage sequence being acted upon by a protease to release BP from the fusion protein.

[0456] In some cases, incorporating a cleavage sequence into BPXTEN is designed to allow the release of a BP that becomes active or more active after release from the XTEN peptide. The cleavage sequence is positioned sufficiently close to the BP sequence, typically within 18, 12, 6, or 2 amino acids from the end of the BP sequence, wherein any remaining residues attached to the BP after cleavage do not significantly interfere with the activity of the BP (e.g., binding to a receptor), but rather provide sufficient proximity to a protease capable of cleaving the cleavage sequence. In some embodiments, the cleavage site is a sequence that can be cleaved by endogenous proteases in humans or animals, wherein BPXTEN can be cleaved after administration to humans or animals. In this case, BPXTEN can act as a prodrug or circulating reservoir of the BP. Examples of cleavage sites considered by the present invention include, but are not limited to, polypeptide sequences that can be cleaved by: mammalian endogenous proteases selected from FXIa, FXIIa, kallikrein, FVIIa, FIXa, FXa, FIIa (thrombin), elastase-2, granzyme B, MMP-12, MMP-13, MMP-17, or MMP-20, or non-mammalian proteases such as TEV, enterokinase, and PreScission. TM Protease (rhinovirus 3C protease) and sorting enzyme A. Sequences known to be cleaved by the aforementioned protease are known in the art. Exemplary cleavage sequences and cleavage sites within sequences, as well as sequence variants, are described in Table 7a. For example, thrombin (activated coagulation factor II) acts on the sequence LTPRSLLV (SEQ ID NO:222) {Rawlings ND et al. (2008) Nucleic Acids Res., 36:D320}, which is cleaved after arginine at position 4 in the sequence. Active FIIa is produced by cleaving FII via FXa in the presence of phospholipids and calcium, and is downstream of factor IX in the coagulation pathway. Once activated, its natural role in coagulation is to cleave fibrinogen, which then sequentially initiates clot formation. FIIa activity is tightly controlled and occurs only when coagulation is necessary for proper hemostasis. However, since coagulation is a continuous process in mammals, by incorporating the LTPRSLLV (SEQ ID NO:223) sequence into BPXTEN between BP and the XTEN peptide, the XTEN peptide will be removed from the adjacent BP when physiological coagulation is required, simultaneously activating extrinsic or intrinsic coagulation pathways, thereby releasing BP over time. Similarly, incorporating other sequences that are acted upon by endogenous proteases into BPXTEN will provide sustained release of BP, and in some cases, this can provide a higher degree of activity for BP in the form of a "prodrug" derived from BPXTEN.

[0457] In some cases, only two or three amino acids flanking the cleavage site (four to six amino acids in total) will be incorporated into the cleavage sequence. In other cases, the known cleavage sequence may have one or more deletions or insertions, or one or two or three amino acid substitutions for any one, two, or three amino acids in the known sequence, wherein the deletions, insertions, or substitutions result in decreased or increased sensitivity to the protease, rather than a lack of sensitivity, resulting in the ability to tailor the release rate of BP from XTEN. Exemplary substitutions are shown in Table 7a.

[0458] Table 7a: Protease cleavage sequences used for BP release

[0459]

[0460]

[0461] ↓ indicates the cutting site;

[0462] NA: Not applicable;

[0463] *A list of multiple amino acids before, between, or after the slash indicates alternative amino acids that can replace that position;

[0464] The "-" indicates that any amino acid can replace the corresponding amino acid indicated in the middle column.

[0465] In some embodiments, the BPXTEN fusion protein may include a spacer region sequence, and may also include one or more cleavage sequences configured to release BP from the fusion protein upon action by a protease. In some embodiments, the one or more cleavage sequences may be sequences having at least about 80% (e.g., at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100%) sequence identity with the sequences shown in Table 7a.

[0466] In some embodiments, this disclosure provides a BP release segment peptide (or release segment (RS)) that is a substrate of one or more mammalian proteases associated with or produced by cells found in or near diseased tissue. Such proteases may include, but are not limited to, protease classes such as metalloproteinases, cysteine ​​proteases, aspartic proteases, and serine proteases, including but not limited to the proteases shown in Table 7b. RS can be particularly used for incorporation into human or animal recombinant peptides to confer a prodrug form that can be activated by cleavage of the RS by a mammalian protease. As described herein, RS is incorporated into a human or animal recombinant peptide composition, and the incorporated binding moiety is linked to an XTEN (the configuration of which is described more fully below), wherein upon cleavage of the RS by one or more proteases for which the RS is a substrate, the binding moiety and the XTEN are released from the composition, and the binding moiety, no longer shielded by the XTEN, regains its full potential to bind its ligand. In those recombinant peptide compositions comprising a first antibody fragment and a second antibody fragment, the composition is also referred to herein as an activatable antibody composition (AAC).

[0467] Table 7b: Proteases in target tissues

[0468]

[0469]

[0470]

[0471] In one embodiment, this disclosure provides an activatable recombinant polypeptide comprising a first release region (RS1) sequence, wherein, upon optimal alignment, the first release region sequence has at least 88%, or at least 94%, or 100% sequence identity with a sequence selected from the sequences shown in Table 8a, wherein RS1 is a substrate of one or more mammalian proteases. In other embodiments, this disclosure provides an activatable recombinant polypeptide comprising an RS1 and a second release region (RS2) sequence, wherein, upon optimal alignment, each release region sequence has at least 88%, or at least 94%, or 100% sequence identity with a sequence selected from the sequences shown in Table 8a, wherein RS1 and RS2 are each a substrate of one or more mammalian proteases. In yet another embodiment, the disclosure provides an activatable recombinant polypeptide comprising a first RS (RS1) sequence, wherein, upon optimal alignment, the first RS sequence has at least 90%, at least 93%, at least 97%, or 100% identity with a sequence selected from the sequences shown in Table 8a, wherein RS is a substrate of one or more mammalian proteases. In other embodiments, this disclosure provides an activatable recombinant polypeptide comprising RS1 and a second release segment (RS2) sequence, wherein, upon optimal alignment, each release segment sequence has at least 88%, at least 94%, or 100% sequence identity with a sequence selected from the sequences shown in Table 8b, wherein each of RS1 and RS2 is a substrate of one or more mammalian proteases. In embodiments of the activatable recombinant polypeptide comprising RS1 and RS2, the two release segments may be identical or may have different sequences.

[0472] This disclosure considers release segments that are substrates of one, two, or three different protease classes selected from metalloproteinases, cysteine ​​proteases, aspartic proteases, and serine proteases, including the proteases shown in Table 7b. In one particular feature, the RS acts as a substrate for proteases found or co-localized in close association with diseased tissues or cells, such as, but not limited to, tumors, cancer cells, and inflammatory tissues, and upon cleavage of the RS, releases the binding portion from the composition that is otherwise shielded by the XTEN of the human or animal recombinant polypeptide composition (and therefore has a lower binding affinity for its respective ligands), and regains its full potential to bind target and / or effector cell ligands. In another embodiment, the RS of the human or animal recombinant polypeptide composition comprises an amino acid sequence that is a substrate for a cellular protease localized to the target cell, including but not limited to the proteases shown in Table 7b. In another particular feature of the human or animal recombinant polypeptide composition, the RS, which is a substrate of two or three protease classes, is designed to have sequences capable of being cleaved by different proteases at different positions in the RS sequence. Therefore, RS, which is a substrate of two, three or more protease classes, has two, three or more different cleavage sites in the RS sequence, but cleavage by a single protease still results in the release of the binding moiety and XTEN from the recombinant polypeptide composition containing RS.

[0473] In one embodiment, the RS of this disclosure for incorporation into a human or animal recombinant polypeptide composition is a substrate of one or more proteases, including transmembrane peptidase, enkephalinase (CD10), PSMA, BMP-1, deintegrin and metalloproteinase (ADAM), ADAM8, ADAM9, ADAM10, ADAM12, ADAM15, ADAM17 (TACE), ADAM19, ADAM28 (MDC-L), ADAM with platelet-reactive protein motif (ADAMTS), ADAMTS1, ADAMTS4, ADAMTS5, MMP-1 (collagenase 1), matrix metalloproteinase-1 (MMP-1), and matrix metalloproteinase-1 (MMP-1). Matrix metalloproteinase-1 (MMP-1), matrix metalloproteinase-2 (MMP-2, gelatinase A), matrix metalloproteinase-3 (MMP-3, matrix lysin 1), matrix metalloproteinase-7 (MMP-7, matrix lysin 1), matrix metalloproteinase-8 (MMP-8, collagenase 2), matrix metalloproteinase-9 (MMP-9, gelatinase B), matrix metalloproteinase-10 (MMP-10, matrix lysin 2), matrix metalloproteinase-11 (MMP-11, matrix lysin 3), matrix metalloproteinase-12 (MMP-12, macrophage elastase), matrix metalloproteinase-13 (MMP-13, collagenase 3), matrix metalloproteinase-14 (MM... P-14, MT1-MMP), matrix metalloproteinase-15 (MMP-15, MT2-MMP), matrix metalloproteinase-19 (MMP-19), matrix metalloproteinase-23 (MMP-23, CA-MMP), matrix metalloproteinase-24 (MMP-24, MT5-MMP), matrix metalloproteinase-26 (MMP-26, matrix lysozyme 2), matrix metalloproteinase-27 (MMP-27, CMMP), legumain, cathepsin B, cathepsin C, cathepsin K, cathepsin L, cathepsin S, cathepsin X, cathepsin D, cathepsin E, secretase, urokinase (uP) A) Tissue plasminogen activator (tPA), plasmin, thrombin, prostate-specific antigen (PSA, KLK3), human neutrophil elastase (HNE), elastase, trypsin, transmembrane serine protease type II (TTSP), DESC1, hepsin (HPN), proteolytic enzyme, proteolytic enzyme-2, TMPRSS2, TMPRSS3, TMPRSS4 (CAP2), fibroblast activating protein (FAP), kallikrein-related peptidases (KLK family), KLK4, KLK5, KLK6, KLK7, KLK8, KLK10, KLK11, KLK13, and KLK14. In one embodiment, RS is a substrate of ADAM17. In one embodiment, RS is a substrate of BMP-1.In one embodiment, RS is a substrate of cathepsin. In one embodiment, RS is a substrate of HtrA1. In one embodiment, RS is a substrate of legumain. In one embodiment, RS is a substrate of MMP-1. In one embodiment, RS is a substrate of MMP-2. In one embodiment, RS is a substrate of MMP-7. In one embodiment, RS is a substrate of MMP-9. In one embodiment, RS is a substrate of MMP-11. In one embodiment, RS is a substrate of MMP-14. In one embodiment, RS is a substrate of uPA. In one embodiment, RS is a substrate of proteolytic enzyme. In one embodiment, RS is a substrate of MT-SP1. In one embodiment, RS is a substrate of neutrophil elastase. In one embodiment, RS is a substrate of thrombin. In one embodiment, RS is a substrate of TMPRSS3. In one embodiment, RS is a substrate of TMPRSS4. In one embodiment, the RS of the human or animal recombinant polypeptide composition is a substrate of at least two proteases, said two proteases being legumain, MMP-1, MMP-2, MMP-7, MMP-9, MMP-11, MMP-14, uPA, and a proteolytic enzyme. In another embodiment, the RS of the human or animal recombinant polypeptide composition is a substrate of legumain, MMP-1, MMP-2, MMP-7, MMP-9, MMP-11, MMP-14, uPA, and a proteolytic enzyme.

[0474] Table 8a: BP release segment sequence.

[0475]

[0476]

[0477]

[0478]

[0479]

[0480] Table 8b: Release Segment Sequence

[0481]

[0482]

[0483]

[0484]

[0485]

[0486]

[0487]

[0488]

[0489]

[0490]

[0491]

[0492]

[0493]

[0494]

[0495]

[0496]

[0497]

[0498]

[0499]

[0500]

[0501]

[0502]

[0503]

[0504]

[0505]

[0506] In another aspect, RSs used for incorporation into recombinant peptides in humans or animals can be designed to be selectively sensitive, exhibiting different cleavage rates and efficiencies for various proteases for which they are substrates. Since a given protease can be found at different concentrations in diseased tissues compared to healthy tissues or circulation, including but not limited to tumors, hematologic malignancies, or inflammatory tissues or sites of inflammation, this disclosure provides RSs having individual amino acid sequences modified to have higher or lower cleavage efficiencies for a given protease. This ensures that, upon approach to target cells or tissues and their co-localized proteases, the recombinant peptide preferentially converts from a prodrug form to its active form (i.e., separation and release from the recombinant peptide after RS ​​cleavage via the binding moiety and XTEN) compared to the cleavage rate of the RS in healthy tissues or circulation. The released antibody fragment binding moiety has a greater capacity to bind to ligands in diseased tissues compared to the prodrug form retained in circulation. This selective design improves the therapeutic index of the resulting composition, leading to reduced side effects relative to conventional therapeutic agents that do not incorporate such site-specific activation.

[0507] As used herein, cleavage efficiency is defined as the log2 ratio of the percentage of test substrate containing RS cleaved to the percentage of control substrate AC1611 cleaved in a biochemical assay in which the reaction is carried out (further detailed in the examples), when each is a human or animal protease, wherein the initial substrate concentration is 6 μM, the reaction is incubated at 37°C for 2 hours, and then stopped by the addition of EDTA, wherein the amounts of digested products and uncleaved substrate are analyzed by non-reducing SDS-PAGE to determine the ratio of cleavage percentages. The cleavage efficiency is calculated as follows: Therefore, a cleavage efficiency of -1 means that the amount of test substrate cleaved is 50% compared to the control substrate, while a cleavage efficiency of +1 means that the amount of test substrate cleaved is 200% compared to the control substrate. A higher cleavage rate of the protease relative to the control will result in higher cleavage efficiency, while a slower cleavage rate of the protease relative to the control will result in lower cleavage efficiency. As detailed in the examples, when testing the cleavage rate of individual proteases in in vitro biochemical assays, the control RS sequence AC1611 (RSR-1517) with the amino acid sequence EAGRSANHEPLGLVAT (SEQ ID NO: 8261) was identified as having appropriate baseline cleavage efficiencies for proteases legumain, MMP-2, MMP-7, MMP-9, MMP-14, uPA, and proteolytic enzymes. RS libraries were created by selectively substituting amino acids at various positions in the RS peptides, and evaluation was conducted on a subject group of seven proteases (more fully detailed in the examples), resulting in an overview for establishing guidelines on appropriate amino acid substitutions to achieve RSs with the desired cleavage efficiencies. When preparing RS with the desired cleavage efficiency, substitution with hydrophilic amino acids A, E, G, P, S, and T is preferred; however, other L-amino acids may be substituted at a given position to adjust the cleavage efficiency, provided that the RS retains at least some sensitivity to protease cleavage. Conservative substitution of amino acids in the peptide to preserve or affect activity is entirely within the knowledge and ability of those skilled in the art. In one embodiment, this disclosure provides RS cleaved by a protease selected from legumain, MMP-1, MMP-2, MMP-7, MMP-9, MMP-11, MMP-14, uPA, or a protein lyase, exhibiting a cleavage efficiency at least 0.2 log2, or 0.4 log2, or 0.8 log2, or 1.0 log2 higher in an in vitro biochemical competitive assay compared to the control sequence RSR-1517 having the sequence EAGRSANHEPLGLVAT (SEQ ID NO: 8261) cleaved by the same protease. In another embodiment, this disclosure provides an RS, wherein the RS is cleaved by a protease selected from legumain, MMP-1, MMP-2, MMP-7, MMP-9, MMP-11, MMP-14, uPA, or a protein lyase, and has a cleavage efficiency that is at least 0.2 log2, or 0.4 log2, or 0.8 log2, or 1.0 log2 lower in an in vitro biochemical competitive assay compared to the cleavage of the control sequence RSR-1517 having the sequence EAGRSANHEPLGLVAT (SEQ ID NO: 8261) by the same protease.In one embodiment, this disclosure provides RS, wherein RS is cleaved at a rate of at least 2-fold, at least 4-fold, at least 8-fold, or at least 16-fold by a protease selected from legumain, MMP-1, MMP-2, MMP-7, MMP-9, MMP-11, MMP-14, uPA, or a protein lyase, compared to the control sequence RSR-1517 having the sequence EAGRSANHEPLGLVAT (SEQ ID NO: 8261). In another embodiment, this disclosure provides RS, wherein RS is cleaved at a rate of at most 1 / 2, at most 1 / 4, at most 1 / 8, or at most 1 / 16 by a protease selected from legumain, MMP-1, MMP-2, MMP-7, MMP-9, MMP-11, MMP-14, uPA, or a protein lyase, compared to the control sequence RSR-1517 having the sequence EAGRSANHEPLGLVAT (SEQ ID NO: 8261).

[0508] In another aspect, this disclosure provides an AAC comprising multiple RSs, wherein each RS sequence is selected from the sequence groups shown in Table 8a, and the RSs are linked to each other by 1 to 6 amino acids selected from glycine, serine, alanine, and threonine. In one embodiment, the AAC comprises a first RS and a second RS different from the first RS, wherein each RS sequence is selected from the sequence groups shown in Table 8a, and the RSs are linked to each other by 1 to 6 amino acids selected from glycine, serine, alanine, and threonine. In another embodiment, the AAC comprises a first RS, a second RS different from the first RS, and a third RS different from the first RS and the second RS, wherein each sequence is selected from the sequence groups shown in Table 8a, and the first RS, the second RS, and the third RS are linked to each other by 1 to 6 amino acids selected from glycine, serine, alanine, and threonine. It is particularly contemplated that the multiple RSs of the AAC can be tandemly arranged to form a sequence that can be cleaved by multiple proteases at different cleavage rates or cleavage efficiencies. In another embodiment, this disclosure provides an AAC comprising RS1 and RS2 selected from the sequence groups shown in Tables 8a-8b, and XTEN1 and XTEN2, such as those described above or elsewhere herein, wherein RS1 is fused between XTEN1 and the binding moiety, and RS2 is fused between XTEN2 and the binding moiety. Considering that such compositions would be more readily cleaved by diseased target tissues expressing multiple proteases compared to healthy tissue or in normal circulation, the resulting fragments with the binding moiety would more readily penetrate the target tissue; for example, tumors, and have an enhanced ability to bind to both target cells and effector cells (or, in the case of an AAC designed with a single binding moiety, target cells only) and to link them together.

[0509] The RS disclosed herein can be used as a therapeutic agent included in a recombinant peptide for the treatment of cancer, autoimmune diseases, inflammatory diseases, and other conditions where limited activation of the recombinant peptide is desired. The human or animal composition addresses an unmet need and is superior in one or more respects to conventional antibody therapeutics or bispecific antibody therapeutics that are active after injection, including enhanced terminal half-life, targeted delivery, and improved therapeutic ratio, accompanied by reduced toxicity to healthy tissues.

[0510] In some embodiments, the (fusion) polypeptide comprises a first release segment (RS1) located between the (first) XTEN and the bioactive polypeptide. In some embodiments, the polypeptide further comprises a second release segment (RS2) located between the bioactive polypeptide and the second XTEN. In some embodiments, RS1 and RS2 have identical sequences. In some embodiments, RS1 and RS2 have different sequences. In some embodiments, RS1 comprises an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the sequences shown in Tables 8a-8b. In some embodiments, RS2 comprises an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the sequences shown in Tables 8a-8b. In some embodiments, RS1 and RS2 are each substrates for cleavage by multiple proteases at one, two, or three cleavage sites within each release segment sequence.

[0511] Reference Excerpt

[0512] In some embodiments, the (fusion) polypeptide also includes one or more reference fragments that can be released from the polypeptide upon digestion by a protease. In some embodiments, each of the one or more reference fragments comprises a portion of the bioactive polypeptide. In some embodiments, the one or more reference fragments are single reference fragments that are different in sequence and molecular weight from all other peptide fragments that can be released from the polypeptide upon digestion by a protease.

[0513] peptide mixture

[0514] This document discloses mixtures comprising multiple polypeptides of various lengths; the mixture comprises a first group of polypeptides and a second group of polypeptides. In some embodiments, each polypeptide in the first group comprises a barcode fragment which (a) can be released from the polypeptide by digestion with a protease, and (b) has a sequence and molecular weight different from the sequence and molecular weight of all other fragments that can be released from the first group of polypeptides. In some embodiments, the second group of polypeptides lacks the barcode fragment of the first group of polypeptides. In some embodiments, both the first group of polypeptides and the second group of polypeptides each comprise a reference fragment which (a) is common to both the first group of polypeptides and the second group of polypeptides, and (b) can be released by digestion with a protease. In some embodiments, the ratio of the first group of polypeptides to the polypeptide containing the reference fragment is greater than 0.70. In some embodiments, the ratio of the first group of polypeptides to the polypeptide containing the reference fragment is greater than 0.8, 0.9, 0.95, or 0.98. In some embodiments, the reference fragment appears no more than once in each polypeptide in the first group of polypeptides and the second group of polypeptides. In some embodiments, the protease is a protease that cleaves at the C-terminus of a glutamate residue. In some embodiments, the protease is a Glu-C protease. In some embodiments, the protease is not a trypsin. In some embodiments, peptides of various lengths include peptides comprising at least one extended recombinant peptide (XTEN), such as any peptide described above or anywhere else herein. In some embodiments, the first group of peptides comprises a full-length peptide, wherein the barcode fragment is a portion of the full-length peptide. In some embodiments, the full-length peptide is a (fusion) peptide, such as any peptide described above or anywhere else herein. In some embodiments, the barcode fragment lacks (does not contain) both the N-terminal and C-terminal amino acids of the full-length peptide. In some embodiments, mixtures of peptides of various lengths differ from each other due to N-terminal truncation, C-terminal truncation, or both N-terminal and C-terminal truncation of the full-length peptide. In some embodiments, the first group of peptides and the second group of peptides may differ in one or more pharmacological properties. Non-limiting exemplary properties include...

[0515] Peptide characterization methods

[0516] This document discloses a method for evaluating the relative amounts of a first group of peptides and a second group of peptides in a mixture comprising peptides of various lengths, wherein (1) each peptide in the first group shares a barcode fragment that appears once and only once in the peptide, and (2) each peptide in the second group lacks a barcode fragment shared by the first group of peptides, wherein each peptide in both the first and second groups of peptides comprises a reference fragment. The method may include contacting the mixture with a protease to generate a plurality of proteolytic fragments derived from cleavage of the first and second groups of peptides, wherein the plurality of proteolytic fragments comprises a plurality of reference fragments and a plurality of barcode fragments. The method may further include determining a ratio of the amount of barcode fragments to the amount of reference fragments to evaluate the relative amounts of the first and second groups of peptides. In some embodiments, the barcode fragment appears no more than once in each peptide in the first group of peptides. In some embodiments, the reference fragment appears no more than once in each peptide in both the first and second groups of peptides. In some embodiments, the plurality of proteolytic fragments comprises a plurality of reference fragments and a plurality of barcode fragments. In some embodiments, the protease cleaves a first group of polypeptides and a second group of polypeptides (or polypeptides of various lengths) at the C-terminal side of a glutamate residue, wherein the glutamate residue is subsequently not a proline residue. In some embodiments, the protease is a Glu-C protease. In some embodiments, the protease is not a trypsin. In some embodiments, the step of determining the ratio of the amount of barcode fragment to the amount of reference fragment includes quantifying the barcode fragment and reference fragment from the mixture after the mixture has been contacted with the protease. In some embodiments, the barcode fragment and reference fragment are identified based on their respective masses. In some embodiments, the barcode fragment and reference fragment are identified by mass spectrometry. In some embodiments, the barcode fragment and reference fragment are identified by liquid chromatography-mass spectrometry (LC-MS). In some embodiments, the step of determining the barcode fragment / reference fragment ratio includes isotopically labeled reference fragments and isotopically labeled barcode fragments, or a mixture of both. In some embodiments, polypeptides of various lengths include polypeptides containing at least one extended recombinant polypeptide (XTEN), as described above or anywhere else herein. In some embodiments, XTEN is characterized in that (i) it contains at least 150 amino acids; (ii) at least 90% of the amino acid residues of XTEN are selected from glycine (G), alanine (A), serine (S), threonine (T), glutamate (E), and proline (P); and (iii) it contains at least four different types of amino acids selected from G, A, S, T, E, and P. In some embodiments, when present, a barcode fragment is part of XTEN.In some embodiments, a mixture of peptides of various lengths comprises any peptide as described above or anywhere else herein. In some embodiments, peptides of various lengths comprise full-length peptides and truncated fragments thereof. In some embodiments, peptides of various lengths consist substantially of full-length peptides and truncated fragments thereof. In some embodiments, a mixture of peptides of various lengths differs from one another due to N-terminal truncation, C-terminal truncation, or both N-terminal and C-terminal truncation of the full-length peptide. In some embodiments, the full-length peptide is a peptide as described above or anywhere else herein. In some embodiments, the ratio of the amount of barcode fragment to reference fragment is greater than 0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.98, or 0.99.

[0517] Peptide quantification based on isotype labeling

[0518] In some embodiments, isomeric tagging can be used to determine the ratio of barcode fragments to reference fragments. Those skilled in the art will understand that isomeric tagging is a mass spectrometry strategy used in quantitative proteomics, where a peptide or protein (or a portion thereof) is labeled with various chemical groups that are isomeric (of the same mass) but differ in the distribution of heavy isotopes around their structure. These tags, often referred to as tandem mass tags, are designed such that during tandem mass spectrometry, following high-energy collision-induced dissociation (CID), the mass tag is cleaved at a specific linker region, resulting in reporter ions of different masses. Those skilled in the art will understand that one of the most common isomeric tags is amine reactive tags.

[0519] The enhanced ability to detect and quantify truncated products (e.g., via isotopic labeling) can generate knowledge that can help design manufacturing processes that include purification steps to minimize the presence of unwanted variants in the purified drug substance / product.

[0520] Reorganization of production

[0521] The disclosure herein includes nucleic acids. Nucleic acids may contain polynucleotides (or polynucleotide sequences) encoding (fusion) polypeptides, such as any polypeptides described above or anywhere else herein; or nucleic acids may contain reverse complements of such polynucleotides (or polynucleotide sequences).

[0522] The disclosures herein include expression vectors containing polynucleotide sequences, such as any polynucleotide sequence described in the preceding paragraph, and regulatory sequences operatively linked to the polynucleotide sequences.

[0523] The disclosure herein includes host cells containing, for example, the expression vectors described in the preceding paragraph. In some embodiments, the host cell is a prokaryote. In some embodiments, the host cell is *Escherichia coli*. In some embodiments, the host cell is a mammalian cell.

[0524] In another aspect, this disclosure provides a method for manufacturing human or animal compositions. In one embodiment, the method includes culturing host cells containing a nucleic acid construct encoding a polypeptide or XTEN-containing composition of any of the embodiments described herein, under conditions promoting the expression of a polypeptide or BPXTEN fusion polypeptide, the nucleic acid construct subsequently recovering the polypeptide or BPXTEN fusion polypeptide using a standard purification method (e.g., column chromatography, HPLC, etc.) in which the composition is recovered, wherein at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 99% of the binding fragments of the expressed polypeptide or BPXTEN fusion polypeptide are correctly folded. In another embodiment of the preparation method, the expressed polypeptide or BPXTEN fusion polypeptide is recovered, wherein at least 90%, or at least 95%, or at least 97%, or at least 99% of the polypeptide or BPXTEN fusion polypeptide is recovered in a monomeric, soluble form.

[0525] In another aspect, this disclosure relates to the preparation of peptides and BPXTEN fusion peptides using *E. coli* or mammalian host cells at high fermentation expression levels of functional proteins, and to providing an expression vector encoding constructs usable in the method to produce a composition of cytotoxically active peptide constructs at high expression levels. In one embodiment, the method includes the steps of: 1) preparing a polynucleotide encoding a peptide of any embodiment disclosed herein; 2) cloning the polynucleotide into an expression vector, which may be a plasmid or other vector controlled by appropriate transcription and translation sequences for high-level protein expression in a biological system; 3) transforming a suitable host cell with the expression vector; and 4) culturing the host cell in a conventional nutrient medium under conditions suitable for expression of the peptide composition. The host cell is *E. coli*, if desired. By this method, expression of the polypeptide results in a fermentation titer of the expression product in the host cell of at least 0.05 g / L, or at least 0.1 g / L, or at least 0.2 g / L, or at least 0.3 g / L, or at least 0.5 g / L, or at least 0.6 g / L, or at least 0.7 g / L, or at least 0.8 g / L, or at least 0.9 g / L, or at least 1 g / L, or at least 2 g / L, or at least 3 g / L, or at least 4 g / L, or at least 5 g / L, and wherein at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 99% of the expressed protein is correctly folded. As used herein, the term “correctly folded” means that the antigen-binding fragment component of the composition has the ability to specifically bind its target ligand. In another embodiment, this disclosure provides a method for generating a polypeptide or a BPXTEN fusion polypeptide, the method comprising culturing host cells in a fermentation reaction under conditions of effective expression of the polypeptide product, the host cells containing a vector encoding a polypeptide comprising the polypeptide or the BPXTEN fusion polypeptide, wherein the concentration of the polypeptide product is greater than about 10 mg / g dry weight host cells (mg / g), or at least about 250 mg / g, or about 300 mg / g, or about 350 mg / g, or about 400 mg / g, or about 450 mg / g, or about 500 mg / g of the polypeptide, and wherein the antigen-binding fragment of the expressed protein is correctly folded.In another embodiment, this disclosure provides a method for generating a polypeptide or a BPXTEN fusion polypeptide, the method comprising culturing host cells in a fermentation reaction under conditions of effective expression of the polypeptide product, the host cells comprising a vector encoding a composition, wherein the concentration of the polypeptide product is greater than about 10 mg / g dry weight host cells (mg / g), or at least about 250 mg / g, or about 300 mg / g, or about 350 mg / g, or about 400 mg / g, or about 450 mg / g, or about 500 mg / g of the polypeptide, and wherein the expressed polypeptide product is soluble.

[0526] Pharmaceutical Composition

[0527] This document discloses pharmaceutical compositions comprising a BPXTEN polypeptide, such as any polypeptide described above or anywhere else herein, and one or more pharmaceutically acceptable excipients. In some embodiments, the pharmaceutical compositions are formulated for intradermal, subcutaneous, intravenous, intraarterial, intraperitoneal, intravitreal, intrathecal, or intramuscular administration. In some embodiments, the pharmaceutical compositions are in liquid form. In some embodiments, the pharmaceutical compositions are in a device implanted in the eye or another body site. In some embodiments, the pharmaceutical compositions are in a pre-filled syringe for a single injection. In some embodiments, the pharmaceutical compositions are formulated as a lyophilized powder for reconstitution prior to application.

[0528] In some embodiments, the dosage is administered intradermally, subcutaneously, intravenously, intravitreally (or otherwise injected into the eye), intraarterially, intraperitoneally, intrasheathically, or intramuscularly. In some embodiments, the pharmaceutical composition is administered using a device implanted in the eye or other body site. In some embodiments, the human or animal is a mouse, rat, monkey, or human.

[0529] The pharmaceutical composition can be administered for treatment via any suitable route. Additionally, the pharmaceutical composition may contain other pharmaceutically active compounds or a variety of compounds of the present invention.

[0530] In some embodiments, the pharmaceutical composition may be administered at a therapeutically effective dose. In some of the foregoing cases, the therapeutically effective dose results in an increase in the time spent within the therapeutic window with respect to the fusion protein compared to the corresponding BP of a fusion protein not linked to the fusion protein and administered to a human or animal at a comparable dose.

[0531] In another embodiment, the present invention provides a method for treating a disease, symptom, or condition, comprising administering the pharmaceutical composition to a human or animal using multiple consecutive doses of a pharmaceutical composition, wherein the multiple consecutive doses are administered using a therapeutically effective dosing regimen.

[0532] The BPXTEN polypeptides of the present invention can be formulated according to known methods to prepare pharmaceutically useful compositions, wherein the polypeptides are combined with pharmaceutically acceptable carrier media, such as aqueous solutions or buffers, pharmaceutically acceptable suspensions, and emulsions. Therapeutic formulations for storage are prepared by mixing the active ingredient, having the desired level of purity, with optional physiologically acceptable carriers, excipients, or stabilizers, as described in Remington's Pharmaceutical Sciences, 16th edition, Osol, A. (1980).

[0533] Drug kit

[0534] In another aspect, the present invention provides a kit for facilitating the use of BPXTEN peptides. In one embodiment, the kit comprises, in at least a first container,: (a) a sufficient amount of a BPXTEN fusion protein composition to treat a disease, condition, or ailment upon administration to a person or animal in need; and (b) a pharmaceutically acceptable carrier; together in a formulation prepared for injection or reconstitution with sterile water, buffer, or dextran; along with a label identifying the BPXTEN drug and storage and handling conditions, and leaflets containing: information on the approved indications for the drug, instructions for reconstitution and / or administration of the BPXTEN drug for the approved indications of prevention and / or treatment, appropriate dosage and safety information, and information identifying the batch number and expiration date of the drug. In another embodiment described above, the kit may include a second container carrying a suitable diluent for the BPXTEN composition, providing the user with an appropriate concentration of BPXTEN for delivery to a person or animal.

[0535] Treatment

[0536] This document discloses the use of polypeptides, such as those described above or anywhere else herein, in the preparation of medicaments for treating diseases in humans or animals. In some embodiments, the specific disease to be treated depends on the selection of the bioactive protein. In some embodiments, the disease is cancer.

[0537] This document discloses a method for treating a disease in a person or animal, the method comprising administering to a person or animal in need one or more therapeutically effective doses of a pharmaceutical composition, such as any pharmaceutical composition described above or anywhere else herein. In some embodiments, the disease is cancer. In some embodiments, the pharmaceutical composition is administered to a person or animal as one or more therapeutically effective doses according to a dosing regimen. In some embodiments, the person or animal is a mouse, rat, monkey, or human.

[0538] The following are examples of compositions and composition evaluations of this disclosure. It should be understood that various other embodiments can be practiced in light of the general description provided above.

[0539] Example

[0540] Example 1. XTEN with barcode design using minimal mutation design from generic XTEN

[0541] This example illustrates an exemplary design method for barcoded XTEN peptides by preparing minimal mutations in the amino acid sequence of a generic XTEN peptide (e.g., one of those in Table 3b above). The relevant criteria for performing minimal mutations include one or more of the following: (a) minimizing sequence changes in the corresponding XTEN peptide; (b) minimizing changes in the amino acid composition of the corresponding XTEN peptide; (c) substantially maintaining the net charge of the corresponding XTEN peptide; (d) substantially maintaining the low immunogenicity of the corresponding XTEN peptide; and (e) substantially maintaining the pharmacokinetic properties provided by the XTEN peptide.

[0542] For example, a barcoded XTEN is constructed by performing one or more mutations on the generic XTEN in Table 9, the mutations including the deletion of a glutamate residue, the insertion of a glutamate residue, the substitution of a glutamate residue, or the substitution of a glutamate residue or any combination thereof.

[0543] Table 9. Four general-purpose XTEN peptides for barcode modification

[0544]

[0545]

[0546]

[0547] Example 2. Sequence analysis and selection of barcoded XTEN peptides for fusion with bioactive peptides (“BP”).

[0548] This example illustrates the design and selection of barcoded XTEN peptides (and more than one barcoded XTEN assembled into a group) for fusion with bioactive peptides. Depending on the location of the barcoded fragment within the XTEN and the manner in which the XTEN peptide is fused with a bioactive protein to form a construct containing the XTEN peptide (e.g., XTEN-modified protease-activated T-cell binder (XPAT)), the barcoded fragment can indicate the truncation of the XTEN peptide.

[0549] Two exemplary XTEN peptides (XTEN864 and XTEN288_1) were subjected to in-silico GluC digestion analysis to quantify the peptide fragments that could be released after complete GluC digestion of the XTEN peptides. The in-silico analysis takes into account that for XTEN peptides with consecutive glutamate residues (e.g., “EE”), GluC can cleave after any glutamate residue. As the results summarized in Table 10 below show, the 10-meric peptide sequence “TPGTSTEPSE (SEQ ID NO: 8880)” and the 14-meric peptide sequence “GSAPGSEPATSGSE (SEQ ID NO: 8881)” each appeared once and only once in the longer XTEN864, while all other peptide sequences appeared two or more times in XTEN864. Furthermore, the 14-meric peptide sequence “GSAPGSEPATSGSE (SEQ ID NO: 8881)” also appeared once and only once in the shorter XTEN288_1.

[0550] The uniqueness of candidate barcodes is evaluated relative to all other peptide fragments that can be released from a construct containing an XTEN peptide. Accordingly, a barcode sequence in an XTEN peptide cannot appear anywhere else in a construct containing an XTEN peptide (including any other XTEN peptides contained therein, any bioactive proteins contained therein, or any connections between adjacent components). For example, Table 11 shows a peptide “uniqueness” table for a group containing two XTEN peptides. The 14-meric peptide sequence “GSAPGSEPATSGSE (SEQ ID NO: 8881)” is not unique in the XTEN peptide group containing both XTEN864 and XTEN288 because it is present in both XTEN864 and XTEN288, and therefore cannot be used as a truncated barcode for detecting peptide products containing both XTEN peptides.

[0551] The selection of a barcode (or a set of barcodes) may also involve identifying and determining the appropriate location or position of a candidate barcode within the XTEN peptide. The location or position of the candidate barcode may be related to pharmacological information of the XTEN peptide (and the construct containing the XTEN peptide as a whole), such as truncation of the XTEN peptide beyond a critical length and / or deletions within the XTEN peptide. If XTEN864 is placed at the N-terminus of a product containing the XTEN peptide, and if a 238-amino acid truncation from the N-terminus of the product does not significantly affect the pharmacological properties of the product, then the decapeptide “TPGTSTEPSE (SEQ ID NO: 8880)” may serve as a suitable barcode fragment.

[0552] Table 10. Representative XTEN sequences used for GluC digestion analysis

[0553]

[0554]

[0555] Table 11. Peptide "uniqueness" analysis

[0556]

[0557] All underlined sequences produce a unique GluC peptide.

[0558] Non-XTEN cores are underlined and italicized.

[0559] The barcode peptide is in bold.

[0560] Exemplary barcode peptide sequences are shown in Table 12 below. These barcode sequences should be side-coded according to structural formula (I):

[0561] AAA-Glu-Barcode Peptide-BBB

[0562] Here, "AAA" represents Gly, Ala, Ser, Thr, or Pro, and "BBB" represents Gly, Ala, Ser, or Thr, configured to facilitate efficient release of the barcode peptides via GluC digestion. Notably, inserting each barcode peptide into the XTEN yields an additional unique sequence immediately preceding or following the inserted barcode peptide.

[0563] Table 12. List of suitable barcode peptides

[0564]

[0565] Example 3: Design and selection of XTEN in fully sequenced XTENized peptide constructs

[0566] This example illustrates the design of a complete sequence polypeptide construct containing two XTEN polypeptides, one at the N-terminus and the other at the C-terminus.

[0567] Table 13 below shows the XTEN peptides used in a representative barcoded BPXTEN (containing a barcoded XTEN peptide at both the N-terminus and C-terminus) and a reference BPXTEN (containing a generic XTEN at both the N-terminus and C-terminus). In the representative barcoded BPXTEN, a barcoded XTEN peptide (SEQ ID No. 8014) is fused at the N-terminus of the BP, and another barcoded XTEN peptide (SEQ ID No. 8015) is fused at the C-terminus of the BP. In the reference BPXTEN, a “Ref-N” XTEN peptide (SEQ ID No. 8896) is fused at the N-terminus of the BP, and a “Ref-C” XTEN peptide (SEQ ID No. 8897) is fused at the C-terminus of the BP. The “Ref-N” XTEN polypeptide (SEQ ID No. 8896) is comparable in length to the barcoded XTEN polypeptide SEQ ID No. 8014; and the “Ref-C” XTEN polypeptide (SEQ ID No. 8897) is comparable in length to the barcoded XTEN polypeptide SEQ ID No. 8015. Both the barcoded BPXTEN and the reference BPXTEN contain a reference sequence within the BP component. The reference sequence is unique and differs in molecular weight from all other peptide fragments that can be released from the corresponding BPXTEN after complete digestion by the GluC protease (e.g., according to Example 5). The uniqueness of the reference sequence is evaluated relative to all other peptide fragments that can be released from the BPXTEN construct.

[0568] Table 13. Representative N-terminal and C-terminal XTEN sets used in full-length BPXTEN constructs

[0569]

[0570]

[0571] Example 4: Recombinant construction and production of barcode-encoded XTEN-based fusion peptides

[0572] Examples 4a-4b illustrate the recombinant construction, production, and purification of full-length peptides containing barcoded XTEN peptides using the methods disclosed herein.

[0573] Example 4a. XTEN-encoded fusion peptide containing barcoded XTEN at the C-terminus

[0574] Expression: A construct encoding an XTEN-encoded fusion polypeptide containing an anti-EpCAM single-stranded variable fragment (scFv) and a barcoded XTEN sequence (SEQ ID NO: 8008) of 864 amino acids at the C-terminus was expressed in the patented *Escherichia coli* strain AmE098 and allocated to the periplasm via an N-terminal secretory leader sequence (MKKNIAFLLASMFVFSIATNAYA-) (SEQ ID NO: 8898), which was cleaved during translocation. Fermentation cultures were grown at 37°C in an animal-free composite medium; the temperature was then reduced to 26°C before phosphate depletion. At harvest, the fermentation broth was centrifuged to allow cell clumps to form. At harvest, the total volume and wet cell weight (WCW; clump / supernatant ratio) were recorded, and cells forming clumps were collected and frozen at -80°C.

[0575] Recovery: The frozen cell clumps were resuspended in a lysis buffer (17.7 mM citric acid, 22.3 mM Na₂HPO₄, 75 mM NaCl, 2 mM EDTA, pH 4.0) targeting 30% wet cell weight. The resuspension was allowed to equilibrate at pH 4, and then homogenized via two passes at 800 ± 50 bar while monitoring the output temperature and maintaining it at 15 ± 5 °C. The pH of the homogenate was confirmed to be within the specified range (pH 4.0 ± 0.2).

[0576] Clarification: To reduce endotoxin and host cell impurities, the homogenate was allowed to undergo overnight (15-20 hours) flocculation at low temperature (10±5℃) and acidic conditions (pH 4.0±0.2). To remove insoluble fractions, the flocculated homogenate was centrifuged at 16,900 RCF for 40 minutes at 2-8℃, and the supernatant was retained. The supernatant was diluted approximately 3-fold with Milli-Q water (MQ) and then adjusted to 7±1 mS / cm with 5M NaCl. To remove nucleic acids, lipids, and endotoxins and to act as a filter aid, the supernatant was adjusted to 0.1% (m / m) diatomaceous earth. To maintain the filter aid suspension, the supernatant was mixed via an impeller and allowed to equilibrate for 30 minutes. A filter string consisting of a depth filter followed by a 0.22 μm filter was assembled and then rinsed with MQ. The supernatant was pumped through the filter string while adjusting the flow rate to maintain a pressure drop of 25±5 psig. To adjust the composite buffer system (based on the ratio of citric acid to Na2HPO4) to the desired range for capture chromatography, the filtrate was adjusted with 500 mM Na2HPO4, where the final Na2HPO4 / citric acid ratio was 9.33:1, and the pH of the buffer filtrate was confirmed to be within the specified range (pH 7.0 ± 0.2).

[0577] purification

[0578] AEX Capture: To separate dimers, aggregates, and large truncated products from monomeric products, and to remove endotoxins and nucleic acids, anion exchange (AEX) chromatography was used to capture negatively charged C-terminal XTEN domains. In this study, AEX1 stationary phase (GE Q Sepharose FF), AEX1 mobile phase A (12.2 mM Na2HPO4, 7.8 mM NaH2PO4, 40 mM NaCl), and AEX1 mobile phase B (12.2 mM Na2HPO4, 7.8 mM NaH2PO4, 500 mM NaCl) were used. The column was equilibrated with AEX1 mobile phase A. Based on the total protein concentration measured by dioctylcarboxylic acid (BCA), the filtrate was loaded onto a column targeting 28 ± 4 g / L resin, decanted with AEX1 mobile phase A, and then washed in one step to 30% B. The bound material was eluted at 20 CV with a gradient of 30% B to 60% B. When A220 is ≥100 mAU above the (local) baseline, fractions are collected in 1 CV aliquots. Analysis is performed on SDS-PAGE and SE-HPLC, and the elution fractions are combined.

[0579] IMAC Intermediate Purification: To ensure C-terminal integrity, immobilized metal affinity chromatography (IMAC) was used to capture the C-terminal multihistidine tag (His(6)(SEQ ID NO:8031)). In this study, the following IMAC stationary phases were used: GE IMAC Sepharose FF, IMAC mobile phase A (18.3 mM Na2HPO4, 1.7 mM NaH2PO4, 500 mM NaCl, 1 mM imidazole), and IMAC mobile phase B (18.3 mM Na2HPO4, 1.7 mM NaH2PO4, 500 mM NaCl, 500 mM imidazole). The column was packed with zinc solution and equilibrated with IMAC mobile phase A. The AEX1 cell was adjusted to pH 7.8 ± 0.1, 50 ± 5 mS / cm (with 5 M NaCl), and 1 mM imidazole, loaded onto an IMAC column targeting 2 g / L of resin, and eluted with IMAC mobile phase A until the absorbance at 280 nm (A280) returned to the (local) baseline. The bound material was eluted in one step to 25% IMAC mobile phase B. IMAC elution was initiated when A280 was ≥10 mAU above the (local) baseline, directed to a vessel pre-doped with sufficient EDTA to reach 2 mM EDTA in 2CV, and terminated once 2CV was collected. The eluent was analyzed by SDS-PAGE.

[0580] Intermediate purification of Protein L: To ensure N-terminal integrity, Protein L was used to capture the κ domain located near the N-terminus of the BPXTEN molecule (especially aEpCAM scFv). In this study, the following Protein L mobile phases were used: stationary phase (GE Capto L), mobile phase A (16.0 mM citrate, 20.0 mM Na₂HPO₄, pH 4.0 ± 0.1), mobile phase B (29.0 mM citrate, 7.0 mM Na₂HPO₄, pH 2.60 ± 0.02), and mobile phase C (3.5 mM citrate, 32.5 mM Na₂HPO₄, 250 mM NaCl, pH 7.0 ± 0.1). The column was equilibrated with mobile phase C. The IMAC eluent was adjusted to pH 7.0 ± 0.1 and 30 ± 3 mS / cm (using 5 M NaCl and MQ) and loaded onto a protein L column targeting 2 g / L resin. It was then eluted with protein L mobile phase C until the absorbance (A280) at 280 nm returned to the (local) baseline. The column was washed with protein L mobile phase A, and protein L mobile phases A and B were used to achieve low-pH elution. The bound material was eluted at approximately pH 3.0 and collected into containers pre-doped with 0.5 M Na2HPO4 for every 10 aliquots. The fractions were analyzed by SDS-PAGE.

[0581] HIC Refining: Hydrophobic interaction chromatography (HIC) was used to separate N-terminal variants (the four residues at the absolute N-terminus are not essential for protein L binding) and bulk conformational variants. In this study, the following HIC phases were used: stationary phase (GE CaptoPhenyl ImpRes), mobile phase A (20 mM histidine, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.1), and mobile phase B (1 M ammonium sulfate, 20 mM histidine, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.1). The column was equilibrated with mobile phase B. The adjusted protein L eluent was loaded onto a HIC column targeting 2 g / L resin and de-eluted with mobile phase B until the absorbance (A280) at 280 nm returned to the (local) baseline. The column was washed with 50% B. The combined material was eluted at 75 CV with a gradient from 50% B to 0% B. Fractions were collected in 1 CV aliquots when A280 was ≥3 mAU above the (local) baseline. Analysis was performed on SE-HPLC and HI-HPLC, and the elution fractions were combined.

[0582] Preparation: To transfer the product to the preparation buffer and achieve the target concentration (0.5 g / L), anion exchange was used again to capture C-terminal XTEN. In this study, the following mobile phases were used: AEX2 stationary phase (GE Q Sepharose FF), AEX2 mobile phase A (20 mM histidine, 40 mM NaCl, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.2), AEX2 mobile phase B (20 mM histidine, 1 M NaCl, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.2), and AEX2 mobile phase C (12.2 mM Na₂HPO₄, 7.8 mM NaH₂PO₄, 40 mM NaCl, 0.02% (w / v) polysorbate 80, pH 7.0 ± 0.2). The column was equilibrated using AEX2 mobile phase C. The HIC cell was adjusted to pH 7.0 ± 0.1 and 7 ± 1 mS / cm (with MQ) and loaded onto an AEX2 column targeting 2 g / L of resin. The column was then eluted with AEX2 mobile phase C until A280 returned to the (local) baseline. The column was washed with AEX2 mobile phase A (20 mM histidine, 40 mM NaCl, 0.02% (w / v) polysorbate 80, pH 6.5 ± 0.2). AEX2 mobile phases A and B were used to generate the {NaCl} step and to achieve elution. The bound material was eluted in one step to 38% AEX2 mobile phase B. AEX2 elution collection was initiated when A280 was ≥ 5 mAU above the (local) baseline and terminated once two column volumes had been collected. The AEX2 eluent was filtered 0.22 μm in a BSC, aliquoted, labeled, and stored as bulk drug substance (BDS) at 80 °C. Bulk active pharmaceutical ingredients (BDS) were confirmed to meet all batch release criteria using various analytical methods. Overall quality was analyzed by SDS-PAGE, monomer / dimer and aggregate ratios were analyzed by SE-HPLC, and N-terminal quality and product homogeneity were analyzed by HI-HPLC.

[0583] Example 4b. An XTEN-modified fusion peptide containing an XTEN barcode at the C-terminus and another XTEN barcode at the N-terminus.

[0584] Expression: A construct encoding an XTEN-encoded fusion polypeptide containing an anti-EGFR single-stranded variable fragment (scFv), an 864-amino acid barcoded XTEN (SEQ ID NO: 8008) at the C-terminus, and a 288-amino acid barcoded XTEN (SEQ ID NO: 8007) at the N-terminus, was expressed in the patented *E. coli* strain AmE098 and allocated to the periplasm via an N-terminal secretory leader sequence (MKKNIAFLLASMFVFSIATNAYA-) (SEQ ID NO: 8898), which was cleaved during translocation. Fermentation cultures were grown at 37°C in an animal-free composite medium; the temperature was then reduced to 26°C before phosphate depletion. At harvest, the fermentation broth was centrifuged to allow cell clumps to form. At harvest, the total volume and wet cell weight (WCW; clump / supernatant ratio) were recorded, and cells forming clumps were collected and frozen at -80°C.

[0585] Recovery: Resuspend the frozen cell clumps in lysis buffer (100 mM citrate) targeting 30% wet cell weight. Allow the resuspension to equilibrate at pH 4.4, then homogenize at 17,000 ± 200 bar while monitoring the output temperature and maintaining it at 15 ± 5 °C. Confirm that the pH of the homogenate is within the specified range (pH 4.4 ± 0.1).

[0586] Clarification: To reduce endotoxin and host cell impurities, the homogenate was allowed to undergo overnight (15-20 hours) flocculation at low temperature (10±5℃) and acidic conditions (pH 4.4±0.1). To remove insoluble fractions, the flocculated homogenate was centrifuged at 8,000 RCF and 2-8℃ for 40 minutes, and the supernatant was retained. To remove nucleic acids, lipids, and endotoxins and to act as a filter aid, the supernatant was adjusted to 0.1% (m / m) diatomaceous earth. To maintain the filter aid suspension, the supernatant was mixed via an impeller and allowed to equilibrate for 30 minutes. A filter string consisting of a depth filter followed by a 0.22 μm filter was assembled and then rinsed with MQ. The supernatant was pumped through the filter string while adjusting the flow rate to maintain a pressure drop of 25±5 psig.

[0587] purification

[0588] Protein L Capture: To remove host cell proteins, endotoxins, and nucleic acids, protein L is used to capture the κ domain present in the aEGFR scFv of the BPXTEN molecule. In this study, a protein L stationary phase (Tosoh TP A...

Claims

1. A method for evaluating the truncation of fusion peptides in a sample, The fusion polypeptide comprises at least one extended recombinant (XTEN) polypeptide fused to the N-terminus or C-terminus of a bioactive polypeptide comprising a reference fragment. The XTEN polypeptide is 150 to 1000 amino acids in length, contains multiple non-overlapping sequence motifs, each non-overlapping sequence motif is 9 to 14 amino acids in length, and at least 90% of the amino acid residues in the XTEN polypeptide are glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), or proline (P). The XTEN comprises a barcode fragment of 4 to 20 amino acids in length, appearing only once in the fusion polypeptide, and wherein: if the XTEN polypeptide is located at the N-terminus of the bioactive polypeptide, the barcode fragment is located within 100 amino acids from the N-terminus of the XTEN polypeptide; if the XTEN polypeptide is located at the C-terminus of the bioactive polypeptide, the barcode fragment is located within 100 amino acids from the C-terminus of the XTEN polypeptide. The sample comprises a first group of peptides and a second group of peptides, which are identical to or truncated from the fusion peptide, wherein each peptide in the first group of peptides retains the barcode fragment, and each peptide in the second group of peptides lacks the barcode fragment, wherein each peptide in both the first group of peptides and the second group of peptides retains the reference fragment, the method comprising: The sample is contacted with a protease to generate a plurality of protease digestion fragments derived from the cleavage of the first group of polypeptides and the second group of polypeptides, wherein the plurality of protease digestion fragments comprise: Multiple reference segments; and Multiple barcode fragments; and The ratio of the amount of the barcode fragment to the amount of the reference fragment is determined to evaluate the relative amounts of the first group of peptides and the second group of peptides, which represent the truncation of the fusion peptide.

2. The method according to claim 1, wherein the XTEN polypeptide has at least 90% sequence identity with a sequence selected from SEQ ID NO: 8001-8019.

3. The method according to claim 1 or 2, wherein the barcode segment comprises a sequence having at least 90% sequence identity with a sequence selected from SEQ ID NO: 8020-8030.

4. The method according to claim 1 or 2, wherein the fusion polypeptide further comprises a second XTEN polypeptide on the other side of the bioactive polypeptide, wherein the second XTEN polypeptide is 150 to 1000 amino acids in length, comprises a plurality of non-overlapping sequence motifs, each non-overlapping sequence motif being 9 to 14 amino acids in length, and at least 90% of the amino acid residues in the XTEN polypeptide are glycine (G), alanine (A), serine (S), threonine (T), glutamic acid (E), or proline (P), and wherein the second XTEN polypeptide comprises a second barcode fragment of 4 to 20 amino acids in length, which appears only once in the fusion polypeptide, and: if the second XTEN polypeptide is located at the C-terminus of the bioactive polypeptide, the second barcode fragment is located within 100 amino acids from the C-terminus of the second XTEN polypeptide; if the second XTEN polypeptide is located at the N-terminus of the bioactive polypeptide, the second barcode fragment is located within 100 amino acids from the N-terminus of the second XTEN polypeptide.

5. The method according to claim 4, wherein the second XTEN polypeptide has at least 90% sequence identity with a sequence selected from SEQ ID NO: 8001-8019.

6. The method of claim 4, wherein the second barcode segment comprises a sequence having at least 90% sequence identity with a sequence selected from SEQ ID NO:8020-8030.

7. The method according to claim 1 or 2, wherein the protease is a Glu-C protease.

8. The method according to claim 1 or 2, wherein the protease is not trypsin.

9. The method of claim 1 or 2, wherein determining the ratio of the amount of barcode fragment to the amount of reference fragment includes quantifying the barcode fragment and the reference fragment from the sample after the sample has been contacted with the protease.

10. The method of claim 9, wherein the barcode segment and the reference segment are identified based on their respective quality.

11. The method of claim 9, wherein the barcode segment and the reference segment are identified by: (i) Mass spectrometry, and / or (ii) Liquid chromatography-mass spectrometry (LC-MS).

12. The method of claim 1 or 2, wherein determining the ratio of the barcode segment to the reference segment comprises: (i) Same quantity different order markers, and / or (ii) The sample is doped with one or both of an isotopically labeled reference fragment and an isotopically labeled barcode fragment.

Citation Information

Patent Citations

  • Treatment or prophylaxis of ischemic heart disease

    US20010027181A1

  • Fibroblast growth factor-19 (FGF-19) nucleic acids and polypeptides and methods of use for the treatment of obesity and related disorders

    US20020042367A1

  • Antibodies that immunospecifically bind to TRAIL receptors

    US20030228309A1

  • Single-Dose Administration of Factor VIIa

    US20080261886A1

  • Extended recombinant polypeptides and compositions comprising same

    US20100239554A1