Stable isotope-labeled peptide library for mass spectrometry-based protein identification
A stable isotope-labeled peptide library addresses the limitations of existing HCP quantitation methods by providing accurate and standardized mass spectrometry-based detection and quantitation, facilitating effective bioprocessing and product quality assurance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-02
AI Technical Summary
Current methods for identifying and quantifying host-cell proteins (HCPs) in biopharmaceuticals, such as enzyme-linked immunosorbent assay (ELISA), are laborious and lack accuracy for specific HCPs, while mass spectrometry-based approaches struggle with consistency and absolute quantitation, limiting their routine application in process development.
A stable isotope-labeled peptide library is developed, comprising sequences from Tables A, B, and C, which serves as internal standards for targeted mass spectrometry, enabling accurate and absolute quantitation of HCPs through parallel-reaction monitoring (PRM) and data-independent acquisition (DIA-MS).
The peptide library provides a comprehensive and standardized method for identifying and quantifying HCPs, enhancing process control strategies by ensuring reliable detection and removal of high-risk HCPs, thereby improving bioprocessing efficiency and product safety.
Smart Images

Figure IMGF000039_0001 
Figure IMGF000040_0001 
Figure IMGF000041_0001
Abstract
Description
Docket No. 0132-0382W01STABLE ISOTOPE-LABELED PEPTIDE LIBRARY FOR MASS SPECTROMETRY-BASED PROTEIN IDENTIFICATIONREFERENCE TO ELECTRONIC SEQUENCE LISTING
[0001] The application contains a Sequence Listing which has been submitted electronically in .XML format and is hereby incorporated by reference in its entirety. Said .XML copy, created on September 24, 2025, is named “0132-0382WO1. xml” and is 192,040 bytes in size. The sequence listing contained in this .XML file is part of the specification and is hereby incorporated by reference herein in its entirety.FIELD OF THE INVENTION
[0002] The present disclosure provides a peptide library comprising stable isotope-labeled (SIL) peptides, wherein the peptides comprise or consist of the sequences according to Table A, Table B, or Table C.BACKGROUND
[0003] Host-cell proteins (HCP) are a source of impurities in biopharmaceuticals, with potential detrimental effects on drug efficacy and safety. Regulatory agencies require impurity testing and enzyme-linked immunosorbent assay (ELISA) is routinely used for total HCP quantitation. However, individual HCP identity remain unknown and monitoring of multiple specific "high risk" HCP require laborious protein-specific assay development. In recent years, liquid chromatography mass spectrometry (LC-MS) has emerged as a complementary approach for HCP analysis, providing identification and quantitation of individual contaminants.SUMMARY OF THE INVENTION
[0004] In some embodiments, the present disclosure provides a peptide library including stable isotope-labeled (SIL) peptides, wherein the peptides include sequences according to Table A. In some embodiments, the present disclosure provides a peptide library including stable isotope-labeled (SIL) peptides, wherein the sequences of the peptides consist of those according to Table A. In some embodiments, the present disclosure provides a peptide library including stable isotope-labeled (SIL) peptides, wherein the peptides include or consist of the sequences according to Table B or Table C.
[0005] In some embodiments, the present disclosure provides an internal standard for mass spectrometry, including at least two SIL peptides of the peptide library. In someDocket No. 0132-0382W01 embodiments, the internal standard includes about 55, about 50, or about 40 SIL peptides of the peptide library. In some embodiments, the present disclosure provides an internal standard for mass spectrometry, including all of the peptides of the peptide library.
[0006] In some embodiments, the present disclosure provides a composition including the internal standard described herein; and a sample including a protein of interest and a host cell protein (HCP). In some embodiments, the protein of interest is an antibody or antigenbinding fragment thereof.
[0007] In some embodiments, the present disclosure provides a method of identifying and / or quantifying a HCP in a sample, comprising performing targeted mass spectrometry on a mixture comprising the sample and the internal standard described herein.
[0008] In some embodiments, the present disclosure provides a method of producing a purified protein of interest, comprising a) expressing the protein of interest in a host cell; b) performing targeted mass spectrometry on a mixture comprising the protein of interest, a HCP, and the internal standard described herein to quantify the HCP; c) when levels of the HCP are higher than an acceptable level, performing a purification step to substantially remove the HCP, thereby producing the purified protein of interest.
[0009] In some embodiments, the method further includes, prior to performing the targeted mass spectrometry, digesting proteins in the mixture with a protease. In some embodiments, the targeted mass spectrometry is parallel-reaction monitoring (PRM). In some embodiments, the sample includes a bioprocessing sample. In some embodiments, the sample includes a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof. In some embodiments, the sample includes a protein of interest, a polysorbate, or both.
[0010] In some embodiments, the HCP is a Chinese hamster ovary (CHO) cell protein. In some embodiments, the HCP includes a protease, a lipase, or both. In some embodiments, the HCP includes one or more proteins according to FIG. 3. In some embodiments, the HCP includes one or more proteins encoded by a gene selected from Anxal, Clu, Ctsa, Ctsb, Ctsd, Ctsz, Dnpep, Dpp8, Erapl, Gapdh, Gm, Gstpl, H671_lgl493, H671_3gl0129, H671_4gl2234, HSPA8, Htral, I79_007438, 179_013245, 179_018355, Itih5, Lambl, Ldha, Lgals3, Lgals3bp, Lgmn, LOCI 00768693, Lpl, Mmpl9, Pcolce, PGK1, Pkm, Plbl2, Plodl, Plod2, Pltp, PPIA, Prdxl, Prdx2, Prep, Psma7, Ran, Rnpep, Tagln2, Timpl, Tkt, TubalA,Docket No. 0132-0382W01 and Vim. In some embodiments, the protein of interest is an antibody or antigen-binding fragment thereof.
[0011] In some embodiments, the present disclosure provides a method of producing a library of peptides, wherein the peptides include sequences according to Table A, Table B, or Table C, including: a) performing untargeted mass spectrometry on at least one bioprocessing sample to identify an initial set of HCP peptides present in the at least one bioprocessing sample; b) removing peptides from the initial set of HCP peptides, the removed peptides including (i) trypsin mis-cleavages; (ii) fewer than 7 or more than 20 amino acids; (iii) a methionine residue; and (iv) high or low hydrophobicity; and c) generating the library of peptides.
[0012] In some embodiments, the untargeted mass spectrometry is data-independent acquisition mass spectrometry (DIA-MS). In some embodiments, the at least one bioprocessing sample includes a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof. In some embodiments, the method further includes labeling the peptides in the library of peptides with a stable isotope. In some embodiments, the peptides in the library of peptides consist of the sequences according to Table A, Table B, or Table C.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] FIGS. 1 A and IB illustrate a method described in embodiments herein. FIG. 1 A shows an overview of the mass spectrometry (MS)-based strategy for targeted and untargeted screening of the in-process samples. FIG. IB shows a scheme depicting the proteomics workflow used for untargeted analysis of in-process samples, generation of stable-isotope labelled peptides library and subsequent targeted analysis.
[0014] FIG. 2 shows the intensity profile of "high-risk" HCPs in in-process samples quantified by DIA-MS across four sample sets described in embodiments herein.
[0015] FIG. 3 shows a heatmap showing residual intensity of 48 HCPs described in embodiments herein, after two different conditions of first polishing chromatography. Residual intensity was calculated as a percentage of protein intensity remaining after polishing when compared to the loaded material.
[0016] FIG. 4 shows the results from quantitation of: A) Phospholipase B-like (PLBL2), B) serine / threonine peptidase Htra 1 (HTRA1) by PRM and DIA-MS in two selected sampleDocket No. 0132-0382W01 sets. Absolute quantity (PRM) is shown in ng / mg, relative quantity (DIA-MS) is shown as Log2 intensity.
[0017] FIG. 5 shows a summary of the in-process samples described herein and the number of proteins detected by DIA-MS in the samples.
[0018] FIG. 6 shows the LLOQ and linearity for two exemplary peptides described in embodiments herein.
[0019] FIG. 7 shows the PrA-HPLC signal intensity of monoclonal antibody (mAb) products (heavy and light chains) across a whole panel of in-process samples. The dotted lines in each plot represent ± 3 standard deviations from average intensity.
[0020] FIG. 8 shows the effects of tested factors (Collison energy, Isolation width, and Gradient length) on tandem MS scan (scan of peptide fragments; "MS2") intensity (panel A), signal-to-noise level (panel B), and number of data points per elution peak (panel C) for the panel of SIL peptides of Table B.
[0021] FIG. 9 shows the average MS2 intensity of detected SIL peptides measured by varying gradient length (panel A), collision energy (panel B), and isolation width (panel C).DETAILED DESCRIPTION
[0022] As used herein, "a" or "an" may mean one or more. As used herein, when used in conjunction with the word "comprising," the words "a" or "an" may mean one or more than one. As used herein, "another" or "a further" may mean at least a second or more.
[0023] Throughout this application, the term "about" is used to indicate that a value includes the inherent variation of error for the method / device being employed to determine the value, or the variation that exists among the study subjects. Typically, the term "about" is meant to encompass approximately or less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20% or higher variability, depending on the situation. In some embodiments, one of skill in the art will understand the level of variability indicated by the term "about," due to the context in which it is used herein. It should also be understood that use of the term "about" also includes the specifically recited value.
[0024] The term "or" is used to mean "and / or," unless explicitly indicated to refer only to alternatives or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and "and / or."Docket No. 0132-0382W01
[0025] As used herein, the terms "comprising" (and any variant or form of comprising, such as "comprise" and "comprises"), "having" (and any variant or form of having, such as "have" and "has"), "including" (and any variant or form of including, such as "includes" and "include") or "containing" (and any variant or form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method, composition, and / or kit of the present disclosure. Furthermore, compositions of the present disclosure can be used to achieve methods of the present disclosure.
[0026] The use of the term "for example" and its corresponding abbreviation "e.g." (whether italicized or not) means that the specific terms recited are representative examples and embodiments of the disclosure that are not intended to be limited to the specific examples referenced or cited unless explicitly stated otherwise.
[0027] As used herein, "between" is a range inclusive of the ends of the range. For example, a number between x and y explicitly includes the numbers x and y, and any numbers that fall within x and y.
[0028] As used herein, "protein," "peptide," or "polypeptide" refer to a polymeric form of amino acids, which can be any length. Proteins can include, e.g., antibodies, structural proteins, enzymes, membrane, membrane-associated, and / or transmembrane proteins, transporters, receptors, signaling proteins, and the like. Proteins and / or peptides of the present disclosure also encompass modified proteins, e.g., conjugated to one or more non-peptide substances such as, e.g., a drug, a targeting moiety, a tag such as a visualization tag, and the like. A protein of interest of the present disclosure can be a therapeutic protein, e.g., used in diagnosis, treatment, and / or prevention of a disease or disorder. In some embodiments, a polysorbate described herein can improve stability of the protein in a pharmaceutical formulation. In some embodiments, the therapeutic protein is an antibody. In some embodiments, the therapeutic protein is an antibody-drug conjugate. In some embodiments, a protein preparation described herein includes a protein, e.g., a therapeutic protein.
[0029] In some embodiments, the protein of interest described herein is present in a bioprocessing sample, i.e., a sample produced by or obtained from a cell. A bioprocessing sample is also referred to herein as a "protein preparation."Docket No. 0132-0382W01
[0030] As used herein in the context of protein preparations, "purification" refers to a process in which one or more substances, e.g., proteins, are isolated from a complex mixture, typically cells, tissues, or organisms. A "purified" protein sample or protein preparation can refer to a sample in which one or more non-water-soluble components of a cell, tissue, or organism (such as, e.g., cell membranes, lipids, aggregated proteins or nucleic acids, and other hydrophobic substances) have been reduced or removed and leaving only the soluble components (such as, e.g., soluble proteins). As used herein, "soluble" can refer to the ability of a substance to dissolve in a certain solvent, e.g., a cell culture medium, a buffer, water or an organic solvent. In the context of proteins, "soluble" can also refer to proteins that do not precipitate and / or aggregate in a certain solvent, e.g., a cell culture medium, a buffer, water, or an organic solvent.
[0031] An exemplary purification process can include: growing a cell culture containing the protein of interest, e.g., a therapeutic protein; separating the cells from the culture media; lysing the cells and separating the lysed cells to generate a cell culture supernatant containing the soluble components and a pellet containing the insoluble components; and subjecting the cell culture supernatant to buffer exchange, pH adjustment, centrifugation, filtration (including, e.g., ultrafiltration and / or diafiltration), chromatography, or any combination thereof to generate a purified protein preparation. In some embodiments, a purified protein preparation of the present disclosure is purified by the process described herein. In some embodiments, a partially purified protein preparation of the present disclosure has been subjected to part of the purification process described herein. For example, a partially purified protein preparation may not have been subjected to all of the buffer exchange, pH adjustment, centrifugation, filtration, and / or chromatography steps used for generating the purified protein preparation. In some embodiments, a cell culture supernatant described herein includes a protein of interest of the present disclosure. In some embodiments, a partially purified protein preparation described herein includes a protein of interest of the present disclosure. In some embodiments, a purified protein preparation described herein includes a protein of interest of the present disclosure.
[0032] The present disclosure provides a simplified data-independent acquisition LC-MS workflow that may be applied to in-process samples (e.g., for production of biotherapeutics), which may be purified at different stages and using different purification conditions. In some embodiments, the peptides identified from the samples are from HCP impurities. In some embodiments, the peptides identified from the samples are used as internal standards, e.g., forDocket No. 0132-0382W01 targeted MS, to detect and / or quantify the HCP impurities in other samples. In some embodiments, the peptides identified from the samples are labeled with stable isotopes to generate stable isotope-labeled (SIL) peptides, wherein the SIL peptides are used to construct a peptide library of internal standards. According to an embodiment of the present disclosure, a data-independent acquisition LC-MS workflow was applied to 23 in-process samples for 3 different biotherapeutics purified at different stages and using different purification conditions. The list of HCP-derived peptides detected in the purified fractions was then used for construction of a comprehensive library consisting of 203 labelled peptides from 119 HCP. 45 peptides from difficult-to-remove HCP were then used as internal standards for absolute quantitation by parallel reaction monitoring targeted MS and characterized for limit of quantitation, variability and linearity. In some embodiments, these internal standards could be combined into specific peptide panels to suits process control strategy requirements and could become a core set of critical reagents for accurate quantitation of multiple known HCP in targeted MS assays.
[0033] Biopharmaceuticals have shown a notable growth in the past decade, with more than 100 biological license applications approved by United States Food and Drug Administration (FDA) in the period 2015-2022 (Martin et al., 2023) and reaching 50% of all new drugs approved in 2022 (Senior, 2023). Remarkably, more than 70% of all recombinant proteins approved for therapeutic use are produced in Chinese Hamster Ovary (CHO) cell lines, making it the most popular mammalian host for the production of different molecular formats, including monoclonal antibodies, antibody fragments, antibody-drug conjugates, Fc fusion-proteins and other protein-based therapeutics (Lyu et al. 2022; Walsh and Walsh, 2022).
[0034] In a typical bioprocessing workflow, the protein of interest is produced upstream in the selected cell line and then purified downstream to remove unwanted impurities, such as host cell proteins (HCP). The purification process often involves multiple chromatographic steps and accounts for a large share of the manufacturing cost, up to 50%-80% of the total (Tuameh et al., 2022). The main reason for such a significant investment resides in the fact that regulatory agencies have designated residual HCP as a critical quality attribute of biotherapeutics, as they could potentially affect safety and efficacy of biotherapeutics (USP39-NF34, <1132> 2016). The current gold standard for measuring HCP clearance in biotherapeutics is enzyme-linked immunosorbent assay (HCP -ELISA), where antibodies raised against host cell proteins are used to monitor total HCP amount. However, in additionDocket No. 0132-0382W01 to quantitation, the use of orthogonal methods to reveal the identity of HCP present in the analyzed samples could provide the information required to drive decision-making in purification process development specifically aiming at reducing risks associated with specific "high risk" residual proteins, also referred to herein as "high-risk" HCPs. These include both manufacturing risks linked to enzymes, like proteases and lipases, which could affect drug efficacy and stability, and patient safety risks associated with potential immunogenic and bioactive proteins (Falkenberg et al., 2018; Valente et al., 2018; Jones et al. 2021). MS-based methods for HCP identification and quantitation (HCP -MS) have been identified as complementary approaches ideally suited to provide the identity information required to monitor the content of specific HCP and a new United States Pharmacopeia chapter aiming to provide general guidance in this new emerging analytical field has been recently proposed and published in the Pharmacopeial Forum for public comment (PF 49(3), <1132.1>, 2023).
[0035] In addition, two recent systematic projects have been instrumental in building knowledge around specific HCP present in existing drugs and in collating information about their potential risks. On one side, Molden et al. (2021) analyzed 29 commercially available drugs, approved by the FDA between 1997 and 2004, all produced in CHO cell lines by 14 different companies, while the BioPhorum Development Group (BPDG), in a joint multicompany effort, created a public database (www.biophorum.com / host-cell-proteins) where all known HCP are classified in four risk categories, reporting identity and associated potential risk (Jones et al., 2021). However, in spite of the undisputed crucial role of MS in HCP identification, questions still persist on whether it could provide an approach to HCP quantitation which is simple and accurate enough to be used on a routine basis to support process development, especially where decision-making is based on the reduction of specific unwanted HCP.
[0036] MS offers two main approaches for detection and quantitation of HCPs, i.e. untargeted and targeted methods. Untargeted methods, such as data-dependent acquisition (DDA), do not require any prior knowledge of the analyte, and are therefore well suited for identification and general profiling of HCP in bioprocessing. Still, due to its stochastic nature, conventional label-free DDA is known for a lack of consistency in detection (Bruderer et al., 2015). Recent advances in the proteomics field led to the development of another untargeted approach, based on data-independent acquisition (DIA). DIA outperforms conventional DDA in terms of sensitivity, reproducibility of protein identification and quantification accuracy, asDocket No. 0132-0382W01 demonstrated on wide range of applications including analysis of HCPs in in-process samples (Bruderer et al., 2015; Krasny et al., 2018; Barkovits et al., 2020; Strasser et al., 2021; Hessmann et al., 2023). Moreover, while early applications of DIA needed time- and costconsuming generation of experimental spectral libraries for data processing, latest developments of so-called library-free search engines now allow a simplified direct processing of acquired DIA data (DIA-MS) (Demichev et al., 2019; Sinitcyn et al., 2021; Hessmann et al., 2023). Relevant to the development of quantitative methods, though, untargeted methods only provide relative quantitation and an MS-based workflow applied as a process control strategy should instead ideally include accurate absolute quantitation of the analyte. While ideally placed for identification of unknown HCP, untargeted methods have some specific constrains regarding quantitation. In fact, MS signal is heavily dependent on the physiochemical properties of the measured analyte (Abdul-Khalek et al., 2023) and untargeted methods are mainly used for differential quantitative analysis, where normalized fold changes in MS signal of each specific protein of origin are compared across different conditions (Peng et al., 2024). Nonetheless, when comparing the absolute abundance of different proteins in a sample, they can only provide an estimate based on either statistical assumptions about ionization efficiency or based on MS response to unrelated proteins used as standards (Silva et al., 2006; Ahrne et al., 2013; Kreimer et al., 2017).
[0037] On the other hand, where the identity of the analyzed HCP is known and labelled internal standards are available, targeted analysis such as multiple- or parallel-reaction monitoring (MRM, PRM) is considered the method of choice for absolute quantitation of proteins by mass spectrometry, including residual HCP at low ppm (E et al., 2023; Ji et al., 2023, <1132.1> USP chapter). Published applications include highly sensitive and reliable absolute 16 quantitation of HCP at low ppm levels (Sook et al., 2023; Ji et al., 2023). Targeted methods typically rely on calibration curves of like-for-like internal peptide standards labelled with stable isotopes for establishing linearity of response across a defined concentrations range, reproducibility, and highest accuracy by correcting for ionization efficiency. Moreover, recent efforts in the standardization of these approaches have made them a promising option for robust and simplified quantitative monitoring of tens to hundreds of different targets in routine testing by mass spectrometry (Walker et al., 2017; Gao et al., 2020; Macklin et al., 2020; Barkovits et al. 2021). However, besides the increasing number of MS-based identifications of residual HCP, the lack of experimental knowledge around which specific peptides are suitable to become reliable stable isotope labelled (SIL) peptideDocket No. 0132-0382W01 standards have limited the applications of targeted MS methods to the monitoring of the most commonly known CHO-derived HCP so far.
[0038] The present disclosure provides an LC-MS-based workflow for differential and / or absolute and relative quantitation of HCPs in in-process samples. In some embodiments, this workflow combines elements of targeted and untargeted analysis which could be used for screening and monitoring of clearance of specific unwanted HCPs, e.g., "high risk" HCPs described herein, to inform downstream process optimization. Building on the internal data and knowledge provided by recent systematic reports on CHO-derived residual HCPs (Molden et al., 2021; Jones et al., 2021), the present disclosure provides, in some embodiments, the first version of a comprehensive library of labelled synthetic peptides that could be used as "off-the-shelf internal standards for absolute quantitation of selected HCPs. The presented collection of standard peptides, also referred to herein as "peptide library," characterized in terms of lower limit of quantitation (LLOQ), technical variability (%CV) and linearity (R2) is the most comprehensive library of HCP peptides available for MS-based impurity testing in CHO-derived bioprocessing samples. In some embodiments, this library helps to streamline the accurate quantitative monitoring of multiple known HCPs in a single targeted MS assay and contributes to the standardization of future process control strategies based on mass spectrometry.
[0039] As discussed in embodiments herein, the peptide library provided herein comprises stable isotope-labeled (SIL) peptides, wherein the peptides comprise sequences according to Table A. In some embodiments, the peptides consist of sequences according to Table A. In some embodiments, at least or about 2, at least or about 3, at least or about 4, at least or about 5, at least or about 10, at least or about 15, at least or about 20, at least or about 25, at least or about 30, at least or about 35, at least or about 40, at least or about 45, at least or about 50, at least or about 55, at least or about 60, at least or about 65, at least or about 70, at least or about 75, at least or about 80, at least or about 85, at least or about 90, at least or about 95, or at least or about 100 SIL peptides of the peptide library are used as an internal standard for mass spectrometry. In some embodiments, the internal standard comprises or consists of about 2 to about 100, or about 10 to about 90, or about 20 to about 80, or about 25 to about 75, or about 30 to about 70, or about 40 to about 65, or about 50 to about 60, or about 50, or about 55 SIL peptides of the peptide library. In some embodiments, the internal standard comprises or consists of all of the SIL peptides according to Table A. In some embodiments, the mass spectrometry is targeted mass spectrometry, e.g., PRM and / or MRM. In someDocket No. 0132-0382W01 embodiments, the peptide sequences are as shown in col. 1 of Table A. In some embodiments, the peptide sequences comprise or consist of SEQ ID NOs: 1-211. In some embodiments, the internal standard includes at least one peptide for each of the proteins as shown in FIG. 3, col. 1. In some embodiments, the internal standard includes at least two peptides for each of the proteins as shown in FIG. 3, col. 2. In some embodiments, an internal standard comprising multiple peptides per protein increases the confidence of detection and likelihood of having a peptide that provides quantifiable signal.
[0040] As discussed in embodiments herein, the peptide library provided herein comprises stable isotope-labeled (SIL) peptides, wherein the peptides comprise or consist of sequences according to Table B. In some embodiments, at least or about 2, at least or about 3, at least or about 4, at least or about 5, at least or about 10, at least or about 15, at least or about 20, at least or about 25, at least or about 30, at least or about 35, at least or about 40, at least or about 45, or about 50 SIL peptides of the peptide library are used as an internal standard for mass spectrometry. In some embodiments, the internal standard comprises or consists of about 2 to about 50, or about 5 to about 45, or about 10 to about 40, or about 15 to about 35, or about 20 to about 30, or about 25 SIL peptides of the peptide library. In some embodiments, the internal standard comprises or consists of all of the SIL peptides according to Table B. In some embodiments, the mass spectrometry is targeted mass spectrometry, e.g., PRM and / or MRM. In some embodiments, the peptide sequences are as shown in cols. 3-5 of Table B. In some embodiments, the internal standard includes at least one peptide or at least two peptides for each of the proteins as shown in col. 2 of Table B. As described in embodiments herein, an internal standard comprising multiple peptides per protein increases the confidence of detection and likelihood of having a peptide that provides quantifiable signal.
[0041] As discussed in embodiments herein, the peptide library provided herein comprises stable isotope-labeled (SIL) peptides, wherein the peptides comprise or consist of sequences according to Table C. In some embodiments, at least or about 2, at least or about 3, at least or about 4, at least or about 5, at least or about 10, at least or about 15, at least or about 20, at least or about 25, at least or about 30, at least or about 35, or about 38 SIL peptides of the peptide library are used as an internal standard for mass spectrometry. In some embodiments, the internal standard comprises or consists of all of the SIL peptides according to Table C. In some embodiments, the mass spectrometry is targeted mass spectrometry, e.g., PRM and / or MRM. In some embodiments, the peptide sequences are as shown in cols. 3-5 of Table C. InDocket No. 0132-0382W01 some embodiments, the internal standard includes at least one peptide or at least two peptides for each of the proteins as shown in col. 2 of Table C. As described in embodiments herein, an internal standard comprising multiple peptides per protein increases the confidence of detection and likelihood of having a peptide that provides quantifiable signal.
[0042] In some embodiments, the present disclosure provides a composition comprising the internal standard described herein and a sample comprising a protein of interest and a HCP. In some embodiments, the sample comprises a bioprocessing sample, e.g., for producing the protein of interest from a host cell. In some embodiments, the host cell is a Chinese Hamster Ovary (CHO) cell.
[0043] In some embodiments, the sample comprises a cell culture supernatant (e.g., from culturing the host cells expressing the protein of interest), a cell lysate (e.g., from lysing the host cells expressing the protein of interest), a partially purified protein preparation, a purified protein preparation, or a combination thereof. In some embodiments, a purified protein preparation is prepared by growing a cell culture containing the protein of interest, e.g., a therapeutic protein; separating the cells from the culture media; lysing the cells and separating the lysed cells to generate a cell culture supernatant containing the soluble components and a pellet containing the insoluble components; and subjecting the cell culture supernatant to buffer exchange, pH adjustment, centrifugation, filtration (including, e.g., ultrafiltration and / or diafiltration), chromatography, or any combination thereof. In some embodiments, a partially purified protein preparation has been subjected to part of the purification process described herein. In some embodiments, a partially purified protein preparation has not been subjected to all of the buffer exchange, pH adjustment, centrifugation, filtration, and / or chromatography steps used for generating the purified protein preparation.
[0044] In some embodiments, the bioprocessing sample, cell culture supernatant, cell lysate, and partially purified or purified protein preparation comprises the protein of interest and the HCP. In some embodiments, the sample comprises the protein of interest, a polysorbate, or both. In some embodiments, the polysorbate is capable of stabilizing the protein of interest, e.g., in a pharmaceutical formulation. In some embodiments, the protein of interest comprises a therapeutic protein. In some embodiments, the protein of interest comprises an antibody or antigen-binding fragment thereof. In some embodiments, the protein of interest comprises a monoclonal antibody. In some embodiments, the HCP is capable of degrading the protein of interest and / or the polysorbate in the sample. In some embodiments, the HCP comprises aDocket No. 0132-0382W01 protease, a lipase, or both. In some embodiments, the HCP comprises a protein as shown in FIG. 3, col. 1. In some embodiments, the HCP comprises one or more proteins encoded by a gene selected from Anxal, Clu, Ctsa, Ctsb, Ctsd, Ctsz, Dnpep, Dpp8, Erapl, Gapdh, Grn, Gstpl, H671_lgl493, H671_3gl0129, H671_4gl2234, HSPA8, Htral, 179 007438, 179 013245, 179 018355, Itih5, Lambl, Ldha, Lgals3, Lgals3bp, Lgmn, LOC100768693, Lpl, Mmpl9, Pcolce, PGK1, Pkm, Plbl2, Plodl, Plod2, Pltp, PPIA, Prdxl, Prdx2, Prep, Psma7, Ran, Rnpep, Tagln2, Timpl, Tkt, TubalA, and Vim. In some embodiments, the HCP comprises a protein as shown in Table A, cols. 2-5. In some embodiments, the HCP comprises a protein as shown in Table B or Table C, col. 2. In some embodiments, the HCP comprises one or more of Clusterin, Carboxypeptidase D, Cathepsin B, Cathepsin D, Cathepsin X, C-X-C motif chemokine, Glutathione reductase, Glutathione S-transferase, 78 kDa glucose-regulated protein, HtrA serine peptidase 1, Lipase (Lip A), Beta-hexosaminidase, Cathepsin LI, Lipoprotein lipase, Matrix metallopeptidase 19, Pyruvate kinase, Group XV phospholipase A2 (Lpla2), Phospholipase B-like, Phospholipase D family member 3, Procollagen-lysine 5-dioxygenase, Palmitoyl-protein thioesterase 1, Peroxiredoxin- 1, Sialate O-acetylesterase, Transforming growth factor beta, Thioredoxin, and Thioredoxin-disulfide reductase. In some embodiments, the HCP is a "high risk" HCP as described herein. In some embodiments, the HCP is a CHO cell protein.
[0045] In other embodiments, the present disclosure provides a composition consisting of or consisting essentially of the internal standard described herein and a sample comprising a protein of interest and a HCP. In embodiments that consist essentially of the recited elements, additional components (e.g., excipients, etc.) that do not materially affect the basic and novel character! stic(s) of the claimed compositions (to provide differential and / or absolute and relative quantitation of HCPs in in-process samples) can be included.
[0046] In some embodiments, the present disclosure provides a method of identifying and / or quantifying a HCP in a sample, comprising performing mass spectrometry on a mixture comprising the sample and the internal standard described herein. In some embodiments, the internal standard comprises about 2 to about 100, or about 10 to about 90, or about 20 to about 80, or about 25 to about 75, or about 30 to about 70, or about 40 to about 65, or about 50 to about 60, or about 50, or about 55 or all of the SIL peptides according to Table A. In some embodiments, the internal standard includes at least one peptide or at least two peptides for each of the proteins as shown in col. 2-5 of Table A. In some embodiments, the internal standard includes at least one peptide or at least two peptides for each of the proteins asDocket No. 0132-0382W01 shown in col. 2 of Table B. In some embodiments, the internal standard includes at least one peptide or at least two peptides for each of the proteins as shown in col. 2 of Table C. In some embodiments, the internal standard includes at least one peptide or at least two peptides for each of the proteins as shown in col. 1 of FIG. 3. In some embodiments, the peptide sequences are as shown in col. 1 of Table A. In some embodiments, the internal standard comprises or consists of all of the SIL peptides according to Table B. In some embodiments, the peptide sequences are as shown in cols. 3-5 of Table B. In some embodiments, the internal standard comprises or consists of all of the SIL peptides according to Table C. In some embodiments, the peptide sequences are as shown in cols. 3-5 of Table C. In some embodiments, the peptide sequences comprise or consist of SEQ ID NOs: 1-211.
[0047] In some embodiments, the mass spectrometry is targeted mass spectrometry. In some embodiments, the targeted mass spectrometry is PRM and / or MRM. In some embodiments, the targeted mass spectrometry is PRM. In some embodiments, the sample comprises a bioprocessing sample, a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof, as described herein. In some embodiments, the sample comprises the protein of interest, a polysorbate, or both. In some embodiments, the protein of interest comprises a therapeutic protein. In some embodiments, the protein of interest comprises an antibody or antigen-binding fragment thereof. In some embodiments, the protein of interest comprises a monoclonal antibody. In some embodiments, the polysorbate is capable of stabilizing the protein of interest, e.g., in a pharmaceutical formulation. In some embodiments, the HCP is capable of degrading the protein of interest and / or the polysorbate in the sample. In some embodiments, the HCP comprises a protease, a lipase, or both. In some embodiments, the HCP comprises a protein as shown in FIG. 3, col. 1. In some embodiments, the HCP comprises one or more proteins encoded by a gene selected from Anxal, Clu, Ctsa, Ctsb, Ctsd, Ctsz, Dnpep, Dpp8, Erapl, Gapdh, Grn, Gstpl, H671_lgl493, H671_3gl0129, H671_4gl2234, HSPA8, Htral, I79_007438, 179_013245, 179_018355, Itih5, Lambl, Ldha, Lgals3, Lgals3bp, Lgmn, LOCI 00768693, Lpl, Mmpl9, Pcolce, PGK1, Pkm, Plbl2, Plodl, Plod2, Pltp, PPIA, Prdxl, Prdx2, Prep, Psma7, Ran, Rnpep, Tagln2, Timpl, Tkt, TubalA, and Vim. In some embodiments, the HCP comprises a protein as shown in Table A, cols. 2-5. In some embodiments, the HCP comprises a protein as shown in Table B, col. 2. In some embodiments, the HCP comprises a protein as shown in Table C, col. 2. In some embodiments, the HCP comprises one or more of Clusterin, Carboxypeptidase D, CathepsinDocket No. 0132-0382W01B, Cathepsin D, Cathepsin X, C-X-C motif chemokine, Glutathione reductase, Glutathione S- transferase, 78 kDa glucose-regulated protein, HtrA serine peptidase 1, Lipase (LipA), Betahexosaminidase, Cathepsin LI, Lipoprotein lipase, Matrix metallopeptidase 19, Pyruvate kinase, Group XV phospholipase A2 (Lpla2), Phospholipase B-like, Phospholipase D family member 3, Procollagen-lysine 5-dioxygenase, Palmitoyl-protein thioesterase 1, Peroxiredoxin- 1, Sialate O-acetylesterase, Transforming growth factor beta, Thioredoxin, and Thioredoxin-disulfide reductase. In some embodiments, the HCP is a "high risk" HCP as described herein. In some embodiments, the HCP is a CHO cell protein.
[0048] In some embodiments, the present disclosure provides a method of producing a purified protein of interest, comprising: a) expressing the protein of interest in a host cell; b) performing targeted mass spectrometry on a mixture comprising the protein of interest, a HCP, and the internal standard described herein to quantify the HCP; and c) when levels of the HCP are higher than an acceptable level, performing a purification step to substantially remove the HCP, thereby producing the purified protein of interest.
[0049] In some embodiments, the method further comprises, prior to performing the targeted mass spectrometry, digesting proteins in the mixture with a protease. Exemplary suitable proteases include LysC, trypsin, pepsin, proteinase K, and elastase. In some embodiments, the purification step comprises buffer exchange, pH adjustment, centrifugation, filtration (including, e.g., ultrafiltration and / or diafiltration), chromatography, or combinations thereof. In some embodiments, the targeted mass spectrometry is PRM and / or MRM. In some embodiments, the targeted mass spectrometry is PRM. In some embodiments, the sample comprises a bioprocessing sample, a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof, as described herein. In some embodiments, the sample comprises the protein of interest, a polysorbate, or both. In some embodiments, the protein of interest comprises a therapeutic protein. In some embodiments, the protein of interest comprises an antibody or antigen-binding fragment thereof. In some embodiments, the protein of interest comprises a monoclonal antibody. In some embodiments, the polysorbate is capable of stabilizing the protein of interest, e.g., in a pharmaceutical formulation. In some embodiments, the HCP is capable of degrading the protein of interest and / or the polysorbate in the sample. In some embodiments, the HCP comprises a protease, a lipase, or both. In some embodiments, the HCP comprises a protein as shown in FIG. 3, col. 1. In some embodiments, the HCP comprises one or more proteins encoded by a gene selected fromDocket No. 0132-0382W01Anxal, Clu, Ctsa, Ctsb, Ctsd, Ctsz, Dnpep, Dpp8, Erapl, Gapdh, Grn, Gstpl, H671_lgl493, H671_3gl0129, H671_4gl2234, HSPA8, Htral, I79_007438, 179_013245, 179_018355, Ttih5, Lambl, Ldha, Lgals3, Lgals3bp, Lgmn, LOC100768693, Lpl, Mmpl9, Pcolce, PGK1, Pkm, Plbl2, Plodl, Plod2, Pltp, PPIA, Prdxl, Prdx2, Prep, Psma7, Ran, Rnpep, Tagln2, Timpl, Tkt, TubalA, and Vim. In some embodiments, the HCP comprises a protein as shown in Table A, cols. 2-5. In some embodiments, the HCP comprises a protein as shown in Table B, col. 2. In some embodiments, the HCP comprises a protein as shown in Table C, col. 2. In some embodiments, the HCP comprises one or more of Clusterin, Carboxypeptidase D, Cathepsin B, Cathepsin D, Cathepsin X, C-X-C motif chemokine, Glutathione reductase, Glutathione S -transferase, 78 kDa glucose-regulated protein, HtrA serine peptidase 1, Lipase (Lip A), Beta-hexosaminidase, Cathepsin LI, Lipoprotein lipase, Matrix metallopeptidase 19, Pyruvate kinase, Group XV phospholipase A2 (Lpla2), Phospholipase B-like, Phospholipase D family member 3, Procollagen-lysine 5-di oxygenase, Palmitoyl-protein thioesterase 1, Peroxiredoxin- 1, Sialate O-acetylesterase, Transforming growth factor beta, Thioredoxin, and Thioredoxin-disulfide reductase. In some embodiments, the HCP is a "high risk" HCP as described herein. In some embodiments, the HCP is a CHO cell protein.
[0050] In some embodiments, the present disclosure provides a method of producing a library of peptides, wherein the peptides comprise sequences according to Table A, Table B, or Table C, comprising: a) performing untargeted mass spectrometry on at least one bioprocessing sample to identify an initial set of HCP peptides present in the at least one bioprocessing sample; removing peptides from the initial set of HCP peptides, the removed peptides comprising (i) trypsin mis-cleavages; (ii) fewer than 7 or more than 20 amino acids; (iii) a methionine residue; and (iv) high or low hydrophobicity; and c) generating the library of peptides. In some embodiments, the library of peptides comprise at least one peptide or at least two peptides for each protein as shown in col. 1 of FIG. 3. In some embodiments, the library of peptides comprise at least one peptide or at least two peptides for each protein as shown in col. 2-5 of Table A. In some embodiments, the peptides consist of the sequences according to Table A. In some embodiments, the peptide sequences are as shown in col. 1 of Table A. In some embodiments, the library of peptides comprise at least one peptide or at least two peptides for each protein as shown in col. 2 of Table B. In some embodiments, the peptides consist of the sequences according to Table B. In some embodiments, the peptide sequences are as shown in cols. 3-5 of Table B. In some embodiments, the library of peptides comprise at least one peptide or at least two peptides for each protein as shown in col. 2 ofDocket No. 0132-0382W01Table C. In some embodiments, the peptides consist of the sequences according to Table C. In some embodiments, the peptide sequences are as shown in cols. 3-5 of Table C. In some embodiments, the peptide sequences comprise or consist of SEQ ID NOs: 1-211.
[0051] In some embodiments, the untargeted mass spectrometry is DIA-MS. In some embodiments, the at least one bioprocessing sample comprises a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof, as described herein. In some embodiments, the removed peptides of (iv) have hydrophobicity higher than about 2 or lower than about -1, e.g., according to the Kyte- Doolittle Hydropathy scale. One of ordinary skill in the art would be capable of evaluating hydrophobicity of the peptides using methods known in the field, e.g., as described in Waibi et al., Front Mol Biosci. 9:960194 (2022). In some embodiments, the method further comprises labeling the peptides in the library of peptides with a stable isotope, thereby forming SIL peptides. Methods of stable isotope labeling peptides are known to one of ordinary skill in the art and are described, e.g., in Basu et al., Nat Protoc. 8(1): 1-12 (2011), doi: 10.1038 / nprot.2011.421. In some embodiments, the library of peptides may be used to prepare the internal standard described herein.ReferencesAbdul-Khalek N, Wimmer R, Overgaard MT, Gregersen Echers S. Insight on physicochemical properties governing peptide MSI response in HPLC-ESI-MS / MS: A deep learning approach. Comput Struct Bi otechnol J. 2023 Jul 22;21:3715-3727. doi: 10.1016 / j.csbj.2023.07.027.Ahrne E, Molzahn L, Glatter T, Schmidt A. Critical assessment of proteome-wide label-free absolute abundance estimation strategies. Proteomics. 2013 Sep;13(17):2567-78. doi: 10.1002 / pmic.201300135.Barkovits, K.; Pacharra, S.; Pfeiffer, K.; Steinbach, S.; Eisenacher, M.; Marcus, K.; Uszkoreit, J. Reproducibility, Specificity and Accuracy of Relative Quantification Using Spectral Library-Based Data-independent Acquisition. Mol Cell Proteomics 2020, 19 (1), 181-197. doi: 10.1074 / mcp.ral 19.001714.Barkovits, K., Chen, W., Kohl, M., Bracht, T. (2021). Targeted Protein Quantification Using Parallel Reaction Monitoring (PRM). In: Marcus, K., Eisenacher, M., Sitek, B. (eds) Quantitative Methods in Proteomics. Methods in Molecular Biology, vol 2228. Humana, New York, NY. doi: 10.1007 / 978-l-0716-1024-4_l.Docket No. 0132-0382W01Bruderer, R.; Bernhardt, O. M.; Gandhi, T.; Miladinovic, S. M.; Cheng, L.-Y.; Messner, S.; Ehrenberger, T.; Zanotelli, V.; Butscheid, Y.; Escher, C.; Vitek, O.; Rinner, O.; Reiter, L. Extending the Limits of Quantitative Proteome Profiling with Data-independent Acquisition and Application to Acetaminophen-Treated Three-Dimensional Liver Microtissues. Mol Cell Proteomics 2015, 14 (5), 1400-1410. doi: 10.1074 / mcp.ml 14.044305.Demichev, V.; Messner, C. B.; Vemardis, S. I.; Lilley, K. S.; Raiser, M. DIA-NN: Neural Networks and Interference Correction Enable Deep Proteome Coverage in High Throughput. Nat Methods 2019, 17 (1), 41-44. doi: 10.1038 / s41592-019-0638-x.E SY, Hu Y, Molden R, Qiu H, Li N. Identification and Quantification of a Problematic Host Cell Protein to Support Therapeutic Protein Development. J Pharm Sci. 2023 Mar;112(3):673-679. doi: 10.1016 / j.xphs.2022.10.008.Escher C, Reiter L, MacLean B, Ossola R, Herzog F, Chilton J, MacCoss MJ, Rinner O. Using iRT, a normalized retention time for more targeted measurement of peptides. Proteomics. 2012 Apr; 12(8): 1111-21. doi: 10.1002 / pmic.201100463.Falkenberg H, Waldera-Lupa DM, Vanderlaan M, Schwab T, Krapfenbauer K, Studts JM, Flad T, Waerner T. Mass spectrometric evaluation of upstream and downstream process influences on host cell protein patterns in biopharmaceutical products. Biotechnol Prog. 2019 May;35(3):e2788. doi: 10.1002 / btpr.2788.Gao X, Rawal B, Wang Y, Li X, Wylie D, Liu YH, Breunig L, Driscoll D, Wang F, Richardson DD. Targeted Host Cell Protein Quantification by LC-MRM Enables Biologies Processing and Product Characterization. Anal Chem. 2020 Jan 7;92(l):1007- 1015. doi: 10.1021 / acs.analchem.9b03952.Hessmann, S.; Chery, C.; Sikora, A.-S.; Gervais, A.; Carapito, C. Host Cell Protein Quantification Workflow Using Optimized Standards Combined with Data-independent Acquisition Mass Spectrometry. J Pharm Anal 2023, 13 (5), 494-502. doi: 10.1016 / j.jpha.2023.03.009.Husson G, Delangle A, O'Hara J, Cianferani S, Gervais A, Van Dorsselaer A, Bracewell D, Carapito C. Dual Data-independent Acquisition Approach Combining Global HCP Profiling and Absolute Quantification of Key Impurities during Bioprocess Development. Anal Chem. 2018 Jan 16;90(2): 1241-1247. doi: 10.1021 / acs.analchem.7b03965.Ji Q, Sokolowska I, Cao R, Jiang Y, Mo J, Hu P. A highly sensitive and robust LC-MS platform for host cell protein characterization in biotherapeutics. Biologicals. 2023 May;82: 101675. doi: 10.1016 / j.biologicals.2023.101675.Docket No. 0132-0382W01Jones M, Palackal N, Wang F, Gaza-Bulseco G, Hurkmans K, Zhao Y, Chitikila C, Clavier S, Liu S, Menesale E, Schonenbach NS, Sharma S, Valax P, Waerner T, Zhang L, Connolly T. "High-risk" host cell proteins (HCPs): A multi-company collaborative view. Biotechnol Bioeng. 2021 Aug;l 18(8):2870-2885. doi: 10.1002 / bit.27808.Krasny, L.; Bland, P.; Kogata, N.; Wai, P.; Howard, B. A.; Natrajan, R. C.; Huang, P. H. SWATH Mass Spectrometry as a Tool for Quantitative Profiling of the Matrisome. J Proteomics 2018, 189, 11-22. doi: 10.1016 / j.jprot.2018.02.026.Kreimer S, Gao Y, Ray S, Jin M, Tan Z, Mussa NA, Tao L, Li Z, Ivanov AR, Karger BL. Host Cell Protein Profiling by Targeted and Untargeted Analysis of Data Independent Acquisition Mass Spectrometry Data with Parallel Reaction Monitoring Verification. Anal Chem. 2017 May 16;89(10):5294-5302. doi: 10.1021 / acs.analchem.6b04892.Luo H, Du Q, Qian C, Mlynarczyk M, Pabst TM, Damschroder M, Hunter AK, Wang WK. Formation of transient highly-charged mAb clusters strengthens interactions with host cell proteins and results in poor clearance of host cell proteins by protein A chromatography. J Chromatogr A. 2022 Aug 30; 1679:463385. doi: 10.1016 / j.chroma.2022.463385.Lyu X, Zhao Q, Hui J, Wang T, Lin M, Wang K, Zhang J, Shentu J, Dalby PA, Zhang H, Liu B. The global landscape of approved antibody therapies. Antib Ther. 2022 Sep 6;5(4):233- 257. doi: 10.1093 / abt / tbac021.Macklin A, Khan S, Kislinger T. Recent advances in mass spectrometry based clinical proteomics: applications to cancer research. Clin Proteomics. 2020 May 24; 17: 17. doi: 10.1186 / sl2014-020-09283-w.Martins AC, Albericio F, de la Torre BG. FDA Approvals of Biologies in 2022. Biomedicines. 2023 May 12; 11(5): 1434. doi: 10.3390 / biomedicinesl 1051434.Molden R, Hu M, Yen E S, Saggese D, Reilly J, Mattila J, Qiu H, Chen G, Bak H, Li N. Host cell protein profiling of commercial therapeutic protein drugs as a benchmark for monoclonal antibody -based therapeutic protein development. MAbs. 2021 Jan- Dec;13(l): 1955811. doi: 10.1080 / 19420862.2021.1955811.Nebert DW, Vasiliou V. Analysis of the glutathione S-transferase (GST) gene family. Hum Genomics. 2004 Nov; l(6):460-4. doi: 10.1186 / 1479-7364-1-6-460.Panikulam S, Hanke A, Kroener F, Karie A, Anderka O, Villiger TK, Lebesgue N. Host cell protein networks as a novel co-elution mechanism during protein A chromatography. Biotechnol Bioeng. 2024 May;121(5): 1716-1728. doi: 10.1002 / bit.28678.Docket No. 0132-0382W01Peng H, Wang H, Kong W, Li J, Goh WWB. Optimizing differential expression analysis for proteomics data via high-performing rules and ensemble inference. Nat Commun. 2024 May 9;15(1):3922. doi: 10.1038 / s41467-024-47899-w.Senior M. Fresh from the biotech pipeline: fewer approvals, but biologies gain share. Nat Biotechnol. 2023 Feb;41(2): 174-182. doi: 10.1038 / s41587-022-01630-6.Silva JC, Gorenstein MV, Li GZ, Vissers JP, Geromanos SJ. Absolute quantification of proteins by LCMSE: a virtue of parallel MS acquisition. Mol Cell Proteomics. 2006 Jan;5(l): 144-56. doi: 10.1074 / mcp.M500230-MCP200.Sinitcyn, P.; Hamzeiy, H.; Soto, F. S.; Itzhak, D.; McCarthy, F.; Wichmann, C.; Steger, M.;Ohmayer, U.; Distler, U.; Kaspar-Schoenefeld, S.; Prianichnikov, N.; Yilmaz, Rudolph, J. D.; Tenzer, S.; Perez-Riverol, Y.; Nagaraj, N.; Humphrey, S. J.; Cox, J. MaxDIA Enables Library-Based and Library-Free Data-independent Acquisition Proteomics. Nat Biotechnol 2021, 39 (12), 1563-1573. Doi: 10.1038 / s41587-021-00968- 7.Strasser, L.; Oliviero, G.; Jakes, C.; Zaborowska, I.; Floris, P.; Silva, M. R. d.; Fiissl, F.; Carillo, S.; Bones, J. Detection and Quantitation of Host Cell Proteins in Monoclonal Antibody Drug Products Using Automated Sample Preparation and Data-independent Acquisition LC-MS / MS. J Pharm Anal 2021, 11 (6), 726-731. https: / / doi.Org / 10.1016 / j.jpha.2021.05.002.Tran B, Grosskopf V, Wang X, Yang J, Walker D Jr, Yu C, McDonald P. Investigating interactions between phospholipase B-Like 2 and antibodies during Protein A chromatography. J Chromatogr A. 2016 Mar 18; 1438 :31 -8. doi: 10.1016 / j.chroma.2016.01.047.Tuameh A, Harding SE, Darton NJ. Methods for addressing host cell protein impurities in biopharmaceutical product development. Biotechnol J. 2023 Mar;18(3):e2200115. doi: 10.1002 / biot.202200115.Valente KN, Levy NE, Lee KH, Lenhoff AM. Applications of proteomic methods for CHO host cell protein characterization in biopharmaceutical manufacturing. Curr Opin Biotechnol. 2018 Oct;53: 144-150. doi: 10.1016 / j.copbio.2018.01.004.Valikangas T, Suomi T, Elo LL. A comprehensive evaluation of popular proteomics software workflows for label-free proteome quantification and imputation. Brief Bioinform. 2018 Nov 27;19(6): 1344-1355. doi: 10.1093 / bib / bbx054.Walker DE, Yang F, Carver J, Joe K, Michels DA, Yu XC. A modular and adaptive mass spectrometry -based platform for support of bioprocess development toward optimal hostDocket No. 0132-0382W01 cell protein clearance. MAbs. 2017 May / Jun;9(4):654-663. doi: 10.1080 / 19420862.2017.1303023.Walsh G, Walsh E. Biopharmaceutical benchmarks 2022. Nat Biotechnol. 2022 Dec;40(12): 1722-1760. doi: 10.1038 / s41587-022-01582-x.EMBODIMENTS
[0052] Embodiment 1. A peptide library comprising stable isotope-labeled (SIL) peptides, wherein the peptides comprise sequences according to Table A.
[0053] Embodiment 2. A peptide library comprising stable isotope-labeled (SIL) peptides, wherein the sequences of the peptides consist of those according to Table A.
[0054] Embodiment 3. A peptide library comprising stable isotope-labeled (SIL) peptides, wherein the peptides comprise or consist of the sequences according to Table B or Table C.
[0055] Embodiment 4. An internal standard for mass spectrometry, comprising at least two SIL peptides of the peptide library of any one of Embodiments 1-3.
[0056] Embodiment 5. The internal standard of Embodiment 4, comprising about 55, about 50, or about 40 SIL peptides of the peptide library of any one of Embodiments 1-3.
[0057] Embodiment 6. An internal standard for mass spectrometry, comprising all of the peptides of the peptide library of any one of Embodiments 1-3.
[0058] Embodiment 7. A composition comprising the internal standard of any one of Embodiments 4-6; and a sample comprising a protein of interest and a host cell protein (HCP).
[0059] Embodiment 8. A method of identifying and / or quantifying a HCP in a sample, comprising performing targeted mass spectrometry on a mixture comprising the sample and the internal standard of any one of Embodiments 4-6.
[0060] Embodiment 9. A method of producing a purified protein of interest, comprising: a) expressing the protein of interest in a host cell; b) performing targeted mass spectrometry on a mixture comprising the protein of interest, a HCP, and the internal standard of any one of Embodiments 4-6 to quantify the HCP; c) when levels of the HCP are higher than an acceptable level, performing a purification step to substantially remove the HCP, thereby producing the purified protein of interest.Docket No. 0132-0382W01
[0061] Embodiment 10. The method of Embodiment 9, further comprising, prior to performing the targeted mass spectrometry, digesting proteins in the mixture with a protease.
[0062] Embodiment 11. The method of any one of Embodiments 8-10, wherein the targeted mass spectrometry is PRM.
[0063] Embodiment 12. The composition or method of any one of Embodiments 7-11, wherein the sample comprises a bioprocessing sample.
[0064] Embodiment 13. The composition or method of any one of Embodiments 7-12, wherein the sample comprises a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof.
[0065] Embodiment 14. The composition or method of any one of Embodiments 7-13, wherein the sample comprises a protein of interest, a polysorbate, or both.
[0066] Embodiment 15. The composition or method of any one of Embodiments 7-14, wherein the HCP is a Chinese hamster ovary (CHO) cell protein.
[0067] Embodiment 16. The composition or method of any one of Embodiments 7-15, wherein the HCP comprises a protease, a lipase, or both.
[0068] Embodiment 17. The composition or method of any one of Embodiments 7-16, wherein the HCP comprises one or more proteins encoded by a gene selected from Anxal, Clu, Ctsa, Ctsb, Ctsd, Ctsz, Dnpep, Dpp8, Erapl, Gapdh, Grn, Gstpl, H671_lgl493, H671_3gl0129, H671_4gl2234, HSPA8, Htral, I79_007438, 179_013245, 179_018355, Ttih5, Lambl, Ldha, Lgals3, Lgals3bp, Lgmn, LOC100768693, Lpl, Mmpl9, Pcolce, PGK1, Pkm, Plbl2, Plodl, Plod2, Pltp, PPIA, Prdxl, Prdx2, Prep, Psma7, Ran, Rnpep, Tagln2, Timpl, Tkt, TubalA, and Vim.
[0069] Embodiment 18. The composition or method of any one of Embodiments 7-17, wherein the protein of interest is an antibody or antigen-binding fragment thereof.
[0070] Embodiment 19. A method of producing a library of peptides, wherein the peptides comprise sequences according to Table A, Table B, or Table C, comprising: a) performing untargeted mass spectrometry on at least one bioprocessing sample to identify an initial set of HCP peptides present in the at least one bioprocessing sample; b) removing peptides from the initial set of HCP peptides, the removed peptides comprising (i) trypsin mis-cleavages; (ii) fewer than 7 or more than 20 amino acids; (iii) a methionine residue; and (iv) high or low hydrophobicity; and c) generating the library of peptides.Docket No. 0132-0382W01
[0071] Embodiment 20. The method of Embodiment 19, wherein the untargeted mass spectrometry is data-independent acquisition mass spectrometry (DIA-MS).
[0072] Embodiment 21. The method of Embodiment 19 or 20, wherein the at least one bioprocessing sample comprises a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof.
[0073] Embodiment 22. The method of any one of Embodiments 19-21, further comprising labeling the peptides in the library of peptides with a stable isotope.
[0074] Embodiment 23. The method of any one of Embodiments 19-22, wherein the peptides in the library of peptides consist of the sequences according to Table A.
[0075] Embodiment 24. The method of any one of Embodiments 19-22, wherein the peptides in the library of peptides consist of the sequences according to Table B.
[0076] Embodiment 25. The method of any one of Embodiments 19-22, wherein the peptides in the library of peptides consist of the sequences according to Table C.EXAMPLE 1.Materials and MethodsMonoclonal antibody production and purification
[0077] Overall 4 sample sets for 3 different monoclonal antibodies (mAbs) from IgGl (product B) and IgG4 (product A, C) classes were used in this study. Monoclonal antibodies were produced by a standard expression system and purified by sequential process consisting of Protein A affinity chromatography (Capture) and 2 additional steps (Step 1 and Step 2). Several purification conditions were applied in individual sample sets and combinations of conditions is summarized in FIG. 5. Samples of cell culture supernatant (CCS) containing the product and samples after every purification step were collected for MS analysis as indicated on FIG. 1 A.Monoclonal antibody titer measurement
[0078] Product concentration in the samples was analyzed by analytical protein A affinity chromatography with UV detection at 280 nm. Obtained samples were centrifuged 10 min at 9,300 x g, diluted 2 to 6x in 50mM glycine, 150mM NaCl, pH 8 (binding buffer) and loaded onto Poros Protein A ID cartridge (Perseptive Biosystems) at 2 ml / min for 30 seconds. Captured product was eluted at the same flow rate by switching 100% binding buffer toDocket No. 0132-0382W01100% of 50mM glycine, 150mM NaCl, pH 8, pH 2.5 (elution buffer) and signal intensity at 280 nm was detected by DAD detector. The concentration of the product was calculated using the area under the curve for the samples and mAb standards with known concentration levels.MS sample preparation
[0079] Samples for DIA and PRM analysis were processed by Biognosys according to Biognosys standard protocol which includes sonication in denaturing buffer using Covaris LE220Rsc sonication device followed by reduction and alkylation. Aliquots of each sample that contains equal amount of the product were digested by mixture of LysC (1 :200 protease to protein ratio) and trypsin (1 :50 protease to protein ratio) at 37 °C overnight. Resulting peptides were desalted on Oasis HLB pElution plate according to manufacturer's instructions. Cleaned peptides were dried in SpeedVac. Samples for DIA were dissolved in 1% ACN, 0.1% formic acid containing iRT peptide mix for retention time calibration (Biognosys) and peptide concentration was measured by mBCA assay (Pierce). Samples for PRM measurement were spiked with 55 stable isotope labelled reference peptides at known concentrations. To determine LLOQ and linear range of the PRM assay, reference peptides were spiked in seven 5-fold dilutions (concentration range from 0.008 to 125 fmol / pg) into a matrix composed of pooled fractions of purified human antibody and analyzed in technical triplicates by PRM.LC-MS / MS data acquisition
[0080] DIA data acquisition and raw data processing was performed by Biognosys. For each sample, 3.85 pg of peptides were loaded on a reversed phase 75 pm x 600 mm column (PicoFrit emitter with 10 pm tip from New Objective packed with 1.5pm ReprosilPur Cl 8 particles) and separated using NeoVanquish UHPLC nano-LC system (Thermo Scientific) connected to an Orbitrap Exploris 480 mass spectrometer (Thermo Scientific) equipped with a Nanospray Flex ion source and a FAIMS Pro ion mobility device (Thermo Scientific).). LC solvents were A: water with 0.1 % F A; B : 80 % acetonitrile, 0.1 % FA in water. The nonlinear LC gradient was 1-50% solvent B in 172 minutes followed by a column washing step at 90% B for 5 minutes. Flow rate was set to ramp from 500 to 250 nL / min (min 0: 500 nL / min, min 172: 250 nL / min, washing at 500 nL / min) at 64 °C. The FAIMS DIA method consisted of one full range MSI scan and 34 DIA segments per applied compensation voltage as previously described. PRM data acquisition and raw data processing was performed byDocket No. 0132-0382W01Biognosys. For each sample 1 pg of peptides were loaded to a reverse phased column (PicoFrit emitter with 75 pm inner diameter, 60 cm length and 10 pm tip from New Objective, packed with 1.7pm Charged Surface Hybrid Cl 8 particles from Waters) and separated using EASY-nLC 1200 (Thermo Scientific) connected to Q Exactive HF mass spectrometer (Thermo Scientific) with Nanospray Flex ion source. The LC gradient was 1 - 59% of solvent B in 55 minutes in non-linear increments followed by 90% B for 8 minutes using same solvents as for DIA analysis. Data were acquired in standard DIA mode run for retention time-based scheduling using Biognosys' high-precision iRT concept (Escher et al., 2012). For PRM data acquisition the scheduling window was 8 minutes for each peptide.Data processing
[0081] DIA data were searched against CHO (Uniprot, Cricetulus griseus, 2023-01-01) and human (Swissprot, Homo sapiens 2023-01-01) .fasta databases from UniProt and product sequences using directDIA algorithm in Spectronaut software (Biognosys, version 17.1). False discovery rate threshold was set to 1% on peptide and protein level, N-term acetylation, oxidation (MP), deamidation (NQ) and carbamylation (RK) were set as variable modifications with maximum of 2 missed cleavages allowed. PRM data extraction was carried out using SpectroDive software (Biognosys, version 11) and absolute quantities in fmol / pg of product was calculated based on calibration curve or using single point calibration based on spiked internal standard. Statistical analysis was performed by custom made R scripts using R (version 4.1.2) with R studio (2022.02.2 build 485) and Microsoft Excel.Selection of peptides for PRM library
[0082] All DIA data acquired across all analysed samples was used for selection of peptides included in the PRM library. The dataset was filtered at the protein level to select HCPs of interest. All proteins detected after any Step 2 condition were selected together with any HCP from the Biophorum database that was detected in any of our samples. Any detected protein with known peptidase, esterase or lipase activity (G0:0008233, G0:0016788, G0:0016298) were added. The obtained list of HCPs was filtered at peptide level to select candidate peptides suitable for PRM library using following triaging sequence: The list of peptides for the HCPs of interest was filtered to remove any peptides with trypsin mis-cleavages and peptides shorter than 7 or longer than 20 amino acids were removed. Next, any peptides containing methionine were removed because methionine is prone to oxidation during sample processing. Only peptides with minimal signal intensity of l x 103in CCS or Capture stepDocket No. 0132-0382W01 were included. Peptides with high (>2) or low (< -1) hydrophobicity were removed. Due to possibility of N,Q deamidation, peptides without N or Q in the sequence were preferentially selected if available. Maximum 3 peptides for "high risk" HCPs or 2 peptides for other HCPs of interest were selected.Results and discussionA flexible strategy combining untargeted and targeted methods allowed differential quantitation of residual HCP in a large panel of in-process samples
[0083] The present disclosure and Example provides a flexible MS-based workflow for quantitation of HCPs in in-process samples. The development strategy (FIG. 1A) was based on combination of the targeted method (PRM) for monitoring of absolute quantity of selected HCPs and library -free DIA for untargeted screening during down-stream process optimization. These two methods are connected through the library of standard SIL peptides. The library was built using the data provided by DIA and is utilized by PRM for targeted analysis. This workflow design also allows continuous expanding of the library with new targets obtained from any future screening studies or any other assays developed for specific contaminants such as ligand leaching.
[0084] To establish and test this workflow, a panel of 24 in-process samples were collected for three different monoclonal antibodies at different stages of the purification process (FIG. IB). This panel provided a range of complexities that occur in a real downstream process, from crude culture supernatant (CCS) to highly purified material after a capture step based on Protein A (PrA), followed by two additional steps (Step 1 and Step 2). To further increase the variability of analyzed matrices, different purification strategies were combined and different conditions were tested for some of the steps. Final set of samples included matrices from three different conditions for PrA Capture, two different conditions for Step 1 and five different conditions for Step 2 (FIG. 5). The panel of 24 in-process samples was analyzed by DIA-MS to test whether the proposed approach could be used to i) distinguish purification conditions promoting the reduction of individual "high risk" HCP in screening studies, and to ii) generate a list of candidate proteins suitable for a library of stable isotope-labeled (SIL) peptides relevant to varied conditions, multiple products, and different purification stages
[0085] Number of HCPs detected by the library-free DIA-MS at <1% FDR spanned from more than 5,200 HCPs detected in CCS to less than 30 HCPs identified in after Step 2 (FIG. 5), when searched against the public UniProt database for Cricetulus griseus taxon (seeDocket No. 0132-0382W01Materials and Methods). Number of HCP detected in purified products was similar to previously reported results obtained with DIA-MS using HCP-specific spectral libraries (Husson et al., 2017; Hessmann et al., 2023) and demonstrates the flexibility of library-free DIA-MS to be applied on the whole range of in-process samples.
[0086] Normalization can be an important step for quantitative profiling of label-free proteomics data and performing an accurate differential quantitation of residual HCP across different samples. Various normalization methods for MS-based proteomics have been developed and evaluated to identify the best strategy (Valikangas et al., 2018); however, most of the methods are designed for high-complexity samples with hundreds or thousands of identified proteins. As the identification rate shown earlier indicates, analysis of residual HCP in bioprocessing samples does not necessarily meet the criteria and alternative normalization approaches need to be used. In one embodiment, normalization was based on the starting amount of product measured by PrA-HPLC, assuming that equal amount of digested mAb loaded onto the column should result in comparable mAb signal intensity across samples. Intensity of heavy and light chains in the acquired data were then analyzed to identify any outliers defined as value outside the ± 3 standard deviations from average intensity. No outliers were found by this approach (see FIG. 7); therefore, the whole dataset was considered suitable for relative quantitative analysis.
[0087] In terms of numbers, the initial PrA capture step seemed to be able to reduce the identifications from approx. 50% to up to 85%, depending on the conditions used. However, it was only after Step 1, whatever the conditions or type of separation used, that the numbers fell below 100 of individual residual HCP. The addition of Step 2 seemed to be able to reduce them further, at least for some of the conditions tested, where the numbers decreased to approx. 20-30. In some cases an increase was observed, indicating an enrichment of different subsets of HCPs most likely depending on the conditions and separation chemistries applied. The numbers obtained from this set of in-process samples were similar to the identifications obtained in a previous study carried out on commercial preparations of approved mAbs and mAbs-derived biotherapeutics, spanning from as low as 5 to up to 50 per drug with 79 different HCP identified overall (Molden et al., 2021). Interestingly, out of those 79, more than half were identified in at least four different molecules and 15 in at least nine. Moreover, most of them were also part of the list of common CHO-derived HCP independently compiled by Jones et al. (2021), suggesting that although drug- and process-specific residual HCP do exist, a pool of recurring, difficult-to-purify CHO-derived HCP also exists. InDocket No. 0132-0382W01 agreement with this observation, the dataset obtained in this example also included most of the known CHO-derived HCP described in Jones et al. (2021), notably including proteins belonging to 23 out of the 25 families described as "high risk" HCPs, categorized based on their potential impact into i. immunogenicity, ii. biological function in humans, Hi. drug quality, and iv. formulation. For each family, multiple hits were sometimes found. This is most likely due to multiple proteins with similar function typically occurring in mammalian proteomes possibly combined to the low level of experimental annotation of the sequenced CHO genomes, which resulted in pseudogenes and "protein-like" species being annotated with the same description, e.g. the Glutathione transferase (GST) family was represented in the dataset by fifteen different proteins. This is not unusual and has been described for the GST family also for the human genome in the past (Nebert and Vasiliou, 2004).
[0088] FIG. 2 reports the relative quantitation profile of the identified "high risk" families, i.e. seven for 30 immunogenicity (CLU, PRDX, S100-A4, ANXA5, PK, PLBL2 and GST), three for biological function in 31 humans (CCL, CXCL3, and TGF-beta), eleven for drug quality (GRP78, ENO, HTRA1, MMP19, CATB, 32 CATD, CATL1, CATX, CPD, PPIA, and PDI) and five for formulation (CES1, LIPA, LPLA2, LPL, SIAE ). The specific hits reported are either the actual protein recorded in Jones et al. (2021) or, in the case of multiprotein families, the protein showing the longest persistence in the downstream process in at least one sample set. The data show that all the monitored "high risk" HCPs become undetectable after at least one of the polishing steps tested. These include Clusterin (CLU), an abundant HCP identified in 23 out of 29 commercial drugs (Molden et al., 2021), and the well-known "difficult-to-38 remove" PLBL2 (Tran et al., 2016).
[0089] The differential quantitation analysis extended beyond the "high-risk" HCPs. Heatmap on FIG. 3 shows 40 comparison of residual intensity for HCPs after two tested conditions for Step 1. Rapidly decreased or no signal intensity after the polishing step was found for majority of HCPs in panel, indicating good clearance of those HCPs by both conditions. However, obvious differences in clearance efficiency of specific proteins were observed between the tested conditions. Step 1 vl was more efficient in removal of CTSD, DPP8, PRDX1 and VIM but less efficient in removal of CLU, CTSA, H671_lgl493, PKM and PLBL2 when compared to Step 1 v2. Particularly clearance of VIM shows striking difference. No residual intensity was detected after Step 1 vl while intensity of VIM after Step 1 v2 was increased. Interestingly, increased levels of actin (H671_4gl2234) were found after both tested conditions demonstrating that despite the general trend, in some conditions,Docket No. 0132-0382W01 specific HCPs might be effectively enriched likely due to the specific physiochemical microenvironment (Luo et al., 2022; Panikulam et al., 2024).
[0090] In summary, the specific HCP profiles obtained by each of the tested conditions demonstrate that DIA-MS could be successfully used as a screening tool for data-driven decision making in process development allowing the identification of purification conditions which result in the reduction of specific unwanted HCP. Furthermore, the observed overlapping with other published systematic reports on commonly identified HCP seem to suggest that the use of the described varied conditions has allowed the recreation of purification outcomes representative of conditions used in real bioprocessing samples.Creation of a knowledge- and DIA MS-informed library of synthetic peptides for absolute quantitation of residual HCPs in biopharmaceuticals
[0091] The library-free DIA-MS approach resulted in the identification and relative quantitation of a large number of proteins, representing the largest study where the use of varied conditions, multiple products and sampling at different stages of the process has successfully allowed the recreation of purification conditions promoting the reduction of a large number of individual "high risk" HCPs.
[0092] A library of stable-isotope labelled (SIL) peptides was generated for targeted analysis by parallel reaction monitoring (PRM). Although use of synthetic SIL peptides as internal standards in targeted approaches is considered the most accurate method for protein absolute quantitation by mass spectrometry, the development of such workflows has been hampered by the large amount of experimental work needed upfront. Moreover, studies published so far only reported a limited number of peptides for which a detailed characterisation has been carried out in order to determine critical assay parameters like Lower limit of quantification (LLOQ), technical variability (% CV) and linearity (R2). Therefore, the previously obtained DIA-MS dataset, where large number of proteins where identified, may be a source providing information about specific peptides that could be used in targeted approaches aiming at absolute quantitation of large panels of residual HCPs.
[0093] To identify peptides suitable for the library, the obtained DIA-MS results from all four sample sets were combined, and multi-tier triaging process was performed to select a list of candidates. In the first stage, the triaging process was performed on protein level by selecting specific HCPs of interest detected by DIA-MS. The selected HCPs included all HCPs identified after polishing 2, any detected HCP that is included in the HCP database ofDocket No. 0132-0382W01 common contaminants (biophorum.com / host-cell-proteins / ) and any other detected HCP with lipase, esterase or peptidase activity. Obtained list of HCPs was then triaged at peptide level in the second stage where detected peptides for the selected HCPs were filtered based on sequence length, signal intensity, peptide hydrophobicity, presence of methionine or trypsin mis-cleavage site in the sequence and peptide detection in the late stages of the down-stream process. Detailed parameters for each filter are described in the Materials and methods section. Final library contains 203 peptides representing 119 proteins including 18 out of the 25 HCP families classified as "high-risk" by HCP database. All peptides may be provided as SIL peptide to be used as internal standards. These SIL peptides can be combined into tailored panels to target project- or product-specific HCPs.
[0094] To assess linearity, lower limit of quantification (LLOQ) and technical variability of the SIL peptide library, a panel of 55 SIL peptides (Table A) was selected, representing 23 HCPs detected by DIA-MS in CCS or post Capture step in all 4 sample sets. To mimic buffer and product background in the late stages of the purification process, samples from polishing 2 were pooled and used as a matrix for assay characterization. All 55 selected SIL peptides were mixed and spiked into the pooled blank matrix. Seven calibration points covering range from 0.008 to 125 fmol / pg of product were prepared by serial dilution and analyzed by PRM in triplicates. Linear range was obtained for 45 of the 55 tested SIL peptides with lowest observed LLOQ of 0.04 fmol / pg for peptide GCVTPVK (SEQ ID NO: 212) from Cathepsin LI and GTVVDVR (SEQ ID NO: 188) from Sialate O-acetylesterase corresponding to LLOQ of 1.5 and 2.5 ng / mg of product, respectively (Table A, FIG. 6). Both peptides selected for quantification of Transforming growth factor beta (TGFB1) failed to provide linear range of quantification and therefore Tgfbl was completely removed from the subsequent analysis. Technical variability (% CV) of the 45 SIL peptides with the linear range of quantification was assessed in measurements of technical triplicates. Average CV was 6.97 % ranging from 2.8 % to 12.7% for data above LLOQ. In Table A, "High-risk HCP" column specifies if the protein is considered a "high-risk" HCP as defined by Jones M., et al. 2021, Biotech. Bioeng.PRM provides absolute quantitation of HCPs present in ppm levels
[0095] Absolute quantitation of the 22 proteins represented by 53 SIL peptides characterized in previous part was performed in the same 4 sample sets used for DIA-MS. Mixture of SIL peptides were spiked into each sample at known concentration and used as internal standard. Multiple-point calibration (MPC) was used for the 45 SIL peptides with the characterizedDocket No. 0132-0382W01 linear dynamic range, while the remaining 8 SIL peptides were quantified using single-point calibration (SPC) based on the spike concentration only. The highest concentration of 39,835 ng / mg was measured for Clusterin in CCS from Sample set 2. On the opposite end, lowest concentration above LLOQ was 1.8 ng / mg for Cathepsin LI measured in sample Capture C from Sample set 3.
[0096] Previously obtained DIA-MS data highly correlated with the PRM quantification of the 22 analyzed HCPs (FIG. 3). Both datasets followed the same trend of quantification as shown on FIG. 3 on example of two selected HCPs (PLBL2 and Serine / threonine peptidase Htra 1 - HTRA1). In addition, comparison of quantities obtained by PRM and DIA-MS showed good agreement in fold change differences between individual samples. For instance in sample set 4, concentration of PLBL2 measured by PRM in Step 1 v2 was 9.2-times higher (Log2 fold change of 3.2) compared to Polishing 2 v5, which is in good agreement with the Log2 fold change of 3.44 measured by DIA-MS. Similarly, levels of HTRA1 in the sample set 4 were about 50-times (Log2 fold change ~6) lower in sample after Step 2 v5 step when compared to post-capture samples which correlates with Log2 fold change of ~6 measured by DIA-MS. The correlation of PRM in DIA-MS was weaker in higher concentrations which might be attributed to the lack of linear relationship between the two methods.
[0097] The data obtained in the present study supports previous conclusions that DIA-MS represent a valuable screening method for differential quantitative analysis of residual HCP across a wide range of in-process samples. However, DIA-MS approaches come with an inherent complexity of execution which requires long acquisition time and skilled expertise for reliable protein identification and statistical data analysis resulting in long turnaround time often unsuitable for purification process development screening requirements. The described targeted approaches using well characterized SIL peptides from the "off-the-shelf library of internal standards, represent an simpler and faster approach for high-throughput monitoring of a substantial number of common CHO-derived residual HCP in the same PRM-MS run.
[0098] Described herein are i) the most comprehensive library of well characterized, ready - to-use SIL peptides capable of accurate quantitation of residual HCP in CHO-derived in- process samples in a single PRM run, and ii) a simplified spectral library-free DIA-MS acquisition method, which could be applied to add relevant targets and continue the expansion of the library to include more SIL peptides for common and / or project-specific HCP. The method has been applied to four independent datasets of in-process samples andDocket No. 0132-0382W01 resulted in identification and relative quantitation of over 5000 different proteins across 24 different in-process samples analyzed in parallel. The panel included three different mAbs, a combination of different purification conditions and separations, and sampling at different stages of the process. The present disclosure provides the largest study of this kind published so far, and the high overlap with previous systematic reports on CHO-derived HCP (Jones et al., 2021, Molden etal., 2021) suggests the panel investigated here is representative of conditions widely used in CHO bioprocessing. The good correlation obtained between absolute quantitation by targeted MS and relative quantitation by DIA-MS demonstrate the potential of the presented DIA-MS workflow as a screening tool, able to provide key input for selection of the best purification strategy focusing on reduction of specific HCPs in process development, providing a flexible workflow for decision-making depending on whether unbiased discovery of unknown HCP or monitoring of known HCP is required. In both cases, the methods have demonstrated their ability to provide key input for selection of the best purification strategy, based on reduction of specific unwanted HCP.
[0099] In addition, for the purpose of identifying peptides which could be suitable for use as internal standards in targeted approaches aiming at the accurate absolute quantitative monitoring of abundant and most commonly occurring HCPs, a library of synthetic peptides was generated. In some embodiments, the library includes 203 peptides representing 119 HCPs. As described herein, an initial set of 55 peptides were tested and 45 successfully characterized in terms of lower limit of quantification (LLOQ), technical variability (% C V) and linearity (R2).
[0100] Moreover, the HCP library described herein represents the largest "off-the-shelf set of labelled synthetic peptides suitable for absolute quantitation of selected HCPs. In some embodiments, the peptides of the peptide library could be combined into ad-hoc peptide panels, e.g., comprising at least two of the peptides, to suits required process control strategy of specific products and could represent the initial core set of critical reagents for accurate quantitation of multiple known HCP in biopharmaceuticals by targeted MS. In some embodiments, the peptide library provided herein provides a simpler, robust, and quick turnaround MS-based method suitable for routine HCP testing. In further embodiments, the peptide library provided herein supports high-throughput screenings and decision-making in purification process development.Docket No. 0132-0382W01EXAMPLE 2.
[0101] In this Example, a further panel of SIL peptides ("peptide panel") was characterized.Selection of Peptide Targets
[0102] The panel of peptide targets include the 50 targets representing 26 HCPs as shown in Table B. Most HCPs were represented by two distinct target peptides. In PRM experiments, it is preferred to target multiple peptides per protein to increase confidence of detection and likelihood of finding a peptide that provides quantifiable signal, particularly in low- abundance proteins such as HCP. Notably, Phospholipase b-like 2 (Plbl2) and Pyruvate kinase (Pkm) were represented by three distinct target peptides each, however Lipase A (Lip A), Lipoprotein lipase (Lpl) Peroxiredoxin- 1 (PRDX1) and Transforming growth factor beta (Tgfbi) were represented only by single target peptide that passed criteria for selection of a PRM candidate.Generation of Spectral Library and Calibration Curves for Peptide Panel
[0103] A spectral library, containing information such as target sequence, retention time, m / z values of the precursors and fragments and intensity of the fragments, was generated from DIA analysis of the pooled SIL peptides in Table B. Acquired data were processed by DIA-NN software which enables searching DIA data and allows subsequent export of the spectral library for detected peptides. The resulting spectral library was imported into Skyline software, which was used for all subsequent processing of PRM data.
[0104] Next, calibration curve for the targets was constructed by serial dilution of the pooled SIL peptide mix and spiking it into a sample of digested monoclonal antibody (mAb) to 1) evaluate initial sensitivity and dynamic range in the presence of mAb background 2) determine the appropriate concentration level for subsequent studies.
[0105] Analysis of the calibration curve by the default unscheduled PRM method resulted in detection of 46 out of 50 peptides in the peptide panel. Peptide target 2 for protein Gstpl was not detected potentially due to incorrect setup of the m / z value outside of the scanned range. The remaining three undetected peptides (Pptl peptide 1, Siae peptide 1, and Siae peptide 2) were not included in the spectral library potentially due to omission during the initial DIA- NN search. Peptides Siae peptide 1 and Siae peptide 2 were later added into the library following re-searching new data.Docket No. 0132-0382W01
[0106] For each peptide, lowest detected spike level was identified. Overall 42 out of the 46 detected peptides were detected at a spike level of 4 fmol / pg of product. This concentration was selected for the subsequent study aimed at optimizing LC-MS parameters for the PRM analysis.Optimization of the PRM Method in Design of Experiment (DoE) Study
[0107] Default PRM method was configured using vendor-recommended parameters supplemented with source and flowrate settings previously optimized for DIA-MS during feasibility stage. To further tune the PRM method, a 2-level full factorial design of experiment (DoE) study was conducted to assess effect of five factors on four different responses:
[0108] List of Factors: 1) Ion injection time; 2) Collision energy; 3) Automatic Gain Control (AGC) target in tandem MS scan (scan of peptide fragments; "MS2"); 4) Isolation width; 5) Gradient length.
[0109] List of Responses: a) Number of detected SIL peptides; b) Total intensity at MS2 level; c) Signal-to-noise level; d) Number of datapoints per elution peak.
[0110] The study was performed using the mAb digest sample spiked by the panel of SIL peptides in Table B at 4 fmol / pg of product. All runs were analyzed by the PRM method with parameter values pre-defined by the DoE study. Results are shown in FIG. 8.
[0111] Overall, 43 peptides were detected of which 42 were detected in all runs. Peptide 1 for procollagen-lysine 5-dioxygenase was missed in one of the runs. A statistical model was not obtained to explain effect of the factors on Number of detected peptides due to low data variation. Thus, this response was removed from data analysis.
[0112] As shown in FIG. 8, Collision energy, Isolation width and Gradient length had significant effect on Total MS2 intensity (FIG. 8, panel A). Isolation width had also significant effect on signal-to-noise ratio (FIG. 8, panel B), while Ion injection time showed significant effect on number of data points (FIG. 8, panel C). All observed effects were independent of one another which enabled the use of One-factor-at-a-time (OF AT) approach for subsequent optimization of individual parameters, as shown in FIG. 9.
[0113] Gradient lengths ranging from 20 to 80 min were evaluated in the OF AT analysis. The highest MS2 signal was observed when using 80 min gradient. A weak gradual decrease in MS2 intensity was noted as the gradient length shortened (FIG. 9, panel A). Based on theDocket No. 0132-0382W01 obtained data, Gradient length of 40 min was selected for all subsequent runs. This gradient length retains >90% of MS2 intensity when compared with the 80 min gradient while significantly improving sample throughput.
[0114] Next, collision energy (CE) was optimized to maximize total MS2 intensity. A range from 24 to 30 was tested with the highest intensity observed at CE of 24. As illustrated in FIG. 9, panel B, the total MS2 intensity of the targeted peptides progressively decreases with increasing CE, resulting in reduction to approximately 70% of initial intensity at CE 30 relative to CE 24. Based on these observations, CE 24 was selected as the setting for all subsequent analyses.
[0115] Optimization of Isolation width in the range from 0.6 m / z to 2 m / z resulted in the highest MS2 signal intensity observed at 0.6 m / z width (FIG. 9, panel C). This width is the lowest value supported by the QE MS system therefore narrower windows could not be tested. As shown in FIG. 8, panel B, isolation width also affects signal to noise ratio. At 0.6 m / z, only four peptides exhibited any measurable background signal. Comparative analysis revealed minimal or no improvement in signal -to-noise ratio with wider windows. In fact, number of peptides with detectable background signal was higher when wider isolation windows were used indicating higher likelihood of potential signal interferences. Based on these findings, 0.6 m / z window was selected as the isolation width for PRM analysis.
[0116] The optimization of Ion injection time aimed to collect at least 8 datapoints per elution peak during data acquisition, ensuring reliable quantification. The selected MS2 resolution on QE instrument allows maximum injection time of 50 ms without compromising the scanning speed. Three injection times - 50, 80 and 110 ms - were tested. The ion injection time of 50 ms achieved a median number of datapoints above the threshold value of 8 and was selected as the setting for PRM analysis.
[0117] The selected parameters of the finalized PRM method used in subsequent studies are: Active gradient length: 40 minCollision energy used for all targets: CE 24 Target isolation window: 0.6 m / z MS2 Ion injection time: 50 ms.
[0118] Other parameters:MS1 / MS2 resolution: 35,000 / 17,500MSI Ion injection time: 100 msDocket No. 0132-0382W01Solvent A: 0.1% formic acid in water Solvent B: 0.1% formic acid in CAN Sample load: 100 pg of product.Linearity, Accuracy, and Carryover Test
[0119] Linearity and accuracy were tested using the finalized PRM method. Calibration curve covering 4.5 orders of magnitude for the panel of 50 SIL peptides of Table B and spiked into the digested mAb sample. Only heavy spikes were targeted during the PRM data acquisition. Three repeated acquisitions of the same sample were performed and peptide was considered identified if detected in two out of three repeats.
[0120] A total of 42 peptides were confidently detected, with 25 peptides detected at concentration levels below 10 ng / mg of product. Linear regression analysis of the calibration curves revealed a lack of linearity in the full range of the assay. Inclusion of the top calibration point consistently decreased accuracy, particularly in the lower end of the calibration curve. To address this, the top calibration point was excluded from the regression analysis, resulting in significant improvement in accuracy. 37 of 42 detected peptides showed R2value >0.98 indicating strong linearity. The SIL panel was refined to 39 SIL peptides covering 26 HCPs. 11 out of 26 HCPs were represented by two targets, Plbd2 was represented by three targets, and 14 out of 26 HCPs were represented by single target. The refined panel was characterized using narrowed calibration curve ranging from 0.025 to 25 fmol / pg of product to ensure measurement remained in the linear range of the assay. The "refined panel" is shown in Table C.
[0121] Carryover assessment was tested using the panel of Table B spiked into the mAb digest sample at the highest calibration point (100 fmol / pg of product). The sample was analyzed using the optimized PRM method, followed by two blank runs consisting of injection of loading buffer analyzed by the identical PRM method. This procedure was repeated three times. Carryover was defined as the presence of peptide signal in the first blank exceeding >20% of signal intensity observed at the lowest calibration point measured in previous step. Based on this criterion, carryover was detected for 9 peptides. No carryover was detected in the second blank indicating effective removal of the peptide residuals from the LC / MS system. As a result, inclusion of blank run between samples is recommended to minimize carryover and ensure data integrity.Docket No. 0132-0382W01Robustness and Matrix Effect Test
[0122] Robustness of the signal between 1stand 96thinjection was evaluated using samples of cell culture supernatant (CCS) and Protein A (PrA) eluate, both containing a test product. Both samples were digested along with a mAb, spiked by the panel of Table C at 5 fmol / pg of product and distributed on 96 well plate.
[0123] In CCS samples, 2 failed injections were observed. In one case, signal was extremely low leading to over-estimation of results, while in second case signal was completely missing. Both runs were treated as outliers and removed from analysis without any effect on robustness assessment as injections critical for the experiment were not affected. For the remaining runs, light-to-heavy (L / H) ratio of peak area was calculated for each peptide in each sample. Light version of 4 out 39 targeted peptides was not detected hence L / H ratio could not be calculated and robustness for these peptide was not assessed. All 35 detected peptides passed the success criteria defined as intra-assay precision between 1stand 32ndinjection < 30%. In addition, 33 peptides passed the stretch criteria for intra-assay precision < 20% between 1stand 96thinjection.
[0124] Effect of sample matrix on peak area of the SIL peptides was tested using the same set of samples acquired for robustness test. CCS sample represented complex matrix with high content of protein background while PrA eluate and mAb represented typical samples from the start and end of the purification process. These samples have high level of product- related background. Analysis of raw data for the heavy SIL peptides by ANOVA did not show any significant differences. The results confirm that sample nature has no matrix effect on the spiked internal standard and does not affect quantification of targeted HCPs.Characterization of Refined Panel
[0125] The refined panel of Table C was characterized in terms of limit of detection (LOD), lower limit of quantification (LLOQ), intra- and inter-assay precision and accuracy. The panel of SIL peptides was diluted to create adjusted calibration curve covering range from 0.025 to 25 fmol / pg of product and spiked into digested CCS sample containing test product.
[0126] For intra-assay precision three repeated acquisitions were performed from the same sample using the same LCMS instrument and setup. LLOQ was defined as lowest spike level passing target criteria for both precision and accuracy defined in ATP document. LLOQ success criteria are defined in ng / ml however the method uses fixed amount of product loaded on column which means that concentration in ng / ml is dependent on the titer of theDocket No. 0132-0382W01 sample. For purpose of this evaluation, titer of 1 mg / ml was assumed resulting in ng / ml = ng / mg of product. Recovery used for Accuracy calculation was obtained by using equation for peptide-specific calibration curve and average H / L ratio. LOD for each peptide is reported as a last calibration point with confident identification in 2 out of 3 replicates. Extrapolated LOD is also reported as intensity of blank + 3 SD values.
[0127] LLOQ values obtained from the inter-assay precision tests were added to the values obtained from the intra-assay precision. In total, 18 peptides covering 16 different HCPs have both LLOQ values (intra- and inter-assay) <10 ng / mg of product. Robustness study showed good stability of all targets. Additional 4 HCPs can be quantified with LLOQ > 10 ng / mg of product. LLOQ for these impurities ranges from 11 to 30 ng / mg of product. No targets for matrix metalloprotease 19 (Mmpl9) passed the evaluation criteria for precision and accuracy in this specific study. In summary, 18 peptides covering 16 HCP with LLOQ <10 ng / mg of product passed the success criteria in this study.Table A. List of peptide sequences and the related HCP identity selected for the peptide library.Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Docket No. 0132-0382W01Table B.Docket No. 0132-0382W01Docket No. 0132-0382W01Table C.Docket No. 0132-0382W01
[0128] In Tables A, B, and C, the [CAM] notation immediately following a cysteine ("C") refers to cysteine carbamidomethylation, i.e., the thiol group of the cysteine is carboxymethylated. In some embodiments, carboxymethylation of the thiol group prevents the cysteine from reacting and forming disulfide linkages, which may complicate peptide mass analysis.
Claims
Docket No. 0132-0382W01CLAIMS1. A peptide library comprising stable isotope-labeled (SIL) peptides, wherein the peptides comprise sequences according to Table A.
2. A peptide library comprising stable isotope-labeled (SIL) peptides, wherein the sequences of the peptides consist of those according to Table A.
3. A peptide library comprising stable isotope-labeled (SIL) peptides, wherein the peptides comprise or consist of the sequences according to Table B or Table C.
4. An internal standard for mass spectrometry, comprising at least two SIL peptides of the peptide library of any one of claims 1 to 3.
5. The internal standard of claim 4, comprising about 55, about 50, or about 40 SIL peptides of the peptide library of any one of claims 1 to 3.
6. An internal standard for mass spectrometry, comprising all of the peptides of the peptide library of any one of claims 1 to 3.
7. A composition comprising the internal standard of any one of claims 4 to 6; and a sample comprising a protein of interest and a host cell protein (HCP).
8. A method of identifying and / or quantifying a HCP in a sample, comprising performing targeted mass spectrometry on a mixture comprising the sample and the internal standard of any one of claims 4 to 6.
9. A method of producing a purified protein of interest, comprising: a) expressing the protein of interest in a host cell; b) performing targeted mass spectrometry on a mixture comprising the protein of interest, a HCP, and the internal standard of any one of claims 4 to 6 to quantify the HCP;Docket No. 0132-0382W01 c) when levels of the HCP are higher than an acceptable level, performing a purification step to substantially remove the HCP, thereby producing the purified protein of interest.
10. The method of claim 9, further comprising, prior to performing the targeted mass spectrometry, digesting proteins in the mixture with a protease.
11. The method of any one of claims 8 to 10, wherein the targeted mass spectrometry is PRM.
12. The composition or method of any one of claims 7 to 11, wherein the sample comprises a bioprocessing sample.
13. The composition or method of any one of claims 7 to 12, wherein the sample comprises a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof.
14. The composition or method of any one of claims 7 to 13, wherein the sample comprises a protein of interest, a polysorbate, or both.
15. The composition or method of any one of claims 7 to 14, wherein the HCP is a Chinese hamster ovary (CHO) cell protein.
16. The composition or method of any one of claims 7 to 15, wherein the HCP comprises a protease, a lipase, or both.
17. The composition or method of any one of claims 7 to 16, wherein the HCP comprises one or more proteins encoded by a gene selected from Anxal, Clu, Ctsa, Ctsb, Ctsd, Ctsz, Dnpep, Dpp8, Erapl, Gapdh, Gm, Gstpl, H671_lgl493, H671_3gl0129, H671_4gl2234, HSPA8, Htral, I79_007438, 179_013245, 179_018355, Itih5, Lambl, Ldha, Lgals3, Lgals3bp, Lgmn, LOC100768693, Lpl, Mmpl9, Pcolce, PGK1, Pkm, Plbl2, Plodl, Plod2, Pltp, PPIA, Prdxl, Prdx2, Prep, Psma7, Ran, Rnpep, Tagln2, Timpl, Tkt, TubalA, and Vim.Docket No. 0132-0382W0118. The composition or method of any one of claims 7 to 17, wherein the protein of interest is an antibody or antigen-binding fragment thereof.
19. A method of producing a library of peptides, wherein the peptides comprise sequences according to Table A, Table B. or Table C, comprising: a) performing untargeted mass spectrometry on at least one bioprocessing sample to identify an initial set of HCP peptides present in the at least one bioprocessing sample; b) removing peptides from the initial set of HCP peptides, the removed peptides comprising (i) trypsin mis-cleavages; (ii) fewer than 7 or more than 20 amino acids; (iii) a methionine residue; and (iv) high or low hydrophobicity; and c) generating the library of peptides.
20. The method of claim 19, wherein the untargeted mass spectrometry is data- independent acquisition mass spectrometry (DIA-MS).
21. The method of claim 19 or 20, wherein the at least one bioprocessing sample comprises a cell culture supernatant, a cell lysate, a partially purified protein preparation, a purified protein preparation, or a combination thereof.
22. The method of any one of claims 19 to 21, further comprising labeling the peptides in the library of peptides with a stable isotope.
23. The method of any one of claims 19 to 22, wherein the peptides in the library of peptides consist of the sequences according to Table A.
24. The method of any one of claims 19 to 22, wherein the peptides in the library of peptides consist of the sequences according to Table B.
25. The method of any one of claims 19 to 22, wherein the peptides in the library of peptides consist of the sequences according to Table C.
Citation Information
Patent Citations
Quantitative protein analysis
US20220205054A1