Isobaric standards from synthetic proteins and peptides

Synthetic peptides engineered to differ from a background proteome are used to detect interference and processing artifacts in mass spectrometry-based proteomics, addressing the limitations of existing standards and enhancing quantification accuracy and instrument performance.

WO2026136540A1PCT designated stage Publication Date: 2026-06-25FLAGSHIP PIONEERING INNOVATIONS VII LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FLAGSHIP PIONEERING INNOVATIONS VII LLC
Filing Date
2025-12-17
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Existing standards for protein/peptide quantitation in multiplex proteomics, such as yeast triple knockout and HeLa digest, are unreliable for assessing instrument performance and quantitative accuracy due to issues like lack of complexity, ionization problems, and interference from coisolation of peptide precursors, necessitating improved standards for accurate mass spectrometry measurements.

Method used

The development of synthetic peptides and proteins, engineered to differ by at least one amino acid from a background proteome, which are used to create mixtures for spiking into samples to detect interference and processing artifacts in mass spectrometry-based proteomics, allowing for ground truth quantification ratios and quality control.

Benefits of technology

These synthetic peptides provide accurate detection of interference and processing artifacts, enabling improved quantification and quality control in multiplex proteomics, ensuring reliable instrument performance and quantitative accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025060111_25062026_PF_FP_ABST
    Figure US2025060111_25062026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosed methods relate to a strategy for producing ground truth quantification ratios (both synthetic knockouts and simulated dose response gradients) using synthetic peptides mixed with a human proteome background. These, synthetic, bespoke proteins serve as a built in quality control for sample preparation for proteomics workflows. Additionally, mixtures of synthetic peptides are combined to form assembled synthetic proteins that can be spiked into samples at the start of processing (e.g., trypsin digesting, purification, etc.). In this way, problems in processing methods can be discovered. Sequences can be engineered to inform how well samples are being processed.
Need to check novelty before this filing date? Find Prior Art

Description

ISOBARIC STANDARDS FROM SYNTHETIC PROTEINS AND PEPTIDES CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 735,860 filed December 18, 2024, which is incoiporated herein by reference in their entirety.REFERENCE TO A SEQUENCE LISTING XML

[0002] This application contains a Sequence Listing which has been submitted electronically in XML format. The Sequence Listing XML is herein incorporated by reference in its entirety. Said XML file, generated on December 17th, 2025, is named FLG-093WO_SL.xml and is 26,462 bytes in size.BACKGROUND

[0003] Sample multiplexing through isobaric labeling is a common strategy in bottom-up proteomics experiments. This strategy typically leverages Tandem Mass Tags (TMTs) to simultaneously quantify peptides from different biological samples. Many mass spectrometry -based acquisition strategies can be used for quantification though no ground truth standards exist to properly benchmark these methods.

[0004] Existing standards used for protein / peptide quantitation, such as yeast triple knockout (TKO) (Thermo Fisher Scientific TMT1 Iplex Yeast Digest Standard) and HeLa digest (Thermo Fisher Scientific Pierce™ HeLa Protein Digest Standard) are unreliable for assessing instrument performance and quantitative accuracy. For example, the yeast standard suffers in part from a lack of complexity and the HeLa standard cannot be used for benchmarking interference that arises from coisolation of peptide precursors. Furthermore, methods have used non-human proteins (see Navarrete -Perea et al., “Synthetic Knockout Protein Standard for Evaluating Interference in Tandem Mass Tag-Based Proteomics,” Analytical Chemistry (2024), 96(17)); however, issues arise with the resultant peptide fragments from such proteins as they may not ionize well or they may not be very detectable. Thus, there is a need for improved standards to enable more accurate measurements when performing multiplex proteomics using mass spectrometry.1IPTS / 200244493 1SUMMARY OF THE INVENTION

[0005] Disclosed herein are improved standards for performing multiplex proteomics using mass spectrometry. Disclosed herein are methods for determining interference and / or sample processing artifacts in mass spectrometry tag-based proteomics using one or more synthetic peptides. Disclosed herein are methods for detecting a processing anomaly using one or more synthetic peptides. Additionally disclosed herein are kits comprising one or more synthetic peptides. Additionally disclosed herein is a mixture comprising one or more synthetic peptides, such as between 1 to 20 synthetic peptides each comprising a sequence selected from the group consisting of SEQ ID NOs: 1-27.

[0006] Generally, disclosed herein is a new strategy for producing ground truth quantification ratios - both synthetic knockouts and simulated dose response gradients - by including synthetic, non-conserved peptides into a human proteome background. This strategy also allows for using synthetic, bespoke proteins as a built in quality control for sample preparation for proteomics workflows. In various embodiments, the number and variety of peptides, as well as bespoke concentrations of each synthetic peptide can be chosen to account for differences in ionization and flight. In various embodiments, mixtures of synthetic peptides can be assembled to form an assembled synthetic protein that can then be spiked into a sample at the start of processing (e.g., trypsin digesting, purification, etc.). In this way, problems in processing methods can be discovered. Sequences can be engineered to inform how well samples are being processed.

[0007] Disclosed herein is method for determining interference and / or sample processing artifacts in mass spectrometry -based proteomics, the method comprising: providing a mixture of one or more synthetic peptides and one or more background peptides of a background proteome, wherein the one or more synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; optionally performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents; performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal.

[0008] In various embodiments, the one or more synthetic peptides are between 5 and 45 amino acids in length. In various embodiments, the one or more synthetic peptides comprise between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise2IPTS / 200244493 1between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides.

[0009] Additionally disclosed herein is a method for determining interference and / or sample processing artifacts in mass spectrometry -based proteomics, the method comprising: providing a mixture comprising between 2 and 30 synthetic peptides between 5 and 45 amino acids in length and one or more background peptides of a background proteome, wherein the 2 to 30 synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; optionally performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents; performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal. In various embodiments, the mixture comprises about 20 synthetic peptides. In various embodiments, the one or more synthetic peptides are selected from a protein comprising a sequence sharing at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or 100% identity to any one of SEQ ID NOs: 1-27. In various embodiments, the mixture comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 different synthetic peptides each comprising a protein sequence selected from the group consisting of SEQ ID NOs: 1-27.

[0010] In various embodiments, methods further comprise digesting one or more background proteins of a background proteome using protein digestion reagents to generate the one or more background peptides. In various embodiments, the protein digestion reagents comprise trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin. In various embodiments, the one or more peptide labeling agents comprise tandem mass tags (TMTs). In various embodiments, methods further comprise: prior to generating the mixture, obtaining the one or more background peptides of the background proteome by: lysing a plurality of cells; and extracting the one or more background peptides from the lysed plurality of cells.

[0011] In various embodiments, methods further comprise performing thiol labeling of the one or more background peptides in the lysed plurality of cells, optionally wherein the thiol labeling comprises contacting the one or more background peptides with desthiobiotin iodoacetamide (DBIA). In various embodiments, methods further comprise spiking one or more assembled synthetic proteins into one or more of:(i) a sample of the lysed plurality of cells;3IPTS / 200244493 1(ii) a sample of the thiol-labeled background peptides; and(iii) a sample of the extracted one or more background peptides.

[0012] In various embodiments, methods further comprise spiking one or more assembled synthetic proteins into two or more of:(i) a sample of the lysed plurality of cells;(ii) a sample of the thiol-labeled background peptides; and(iii) a sample of the extracted one or more background peptides.

[0013] In various embodiments, methods further comprise spiking one or more assembled synthetic proteins into each of:(i) a sample of the lysed plurality of cells;(ii) a sample of the thiol-labeled background peptides; and(iii) a sample of the extracted one or more background peptides.

[0014] In various embodiments, an assembled synthetic protein is assembled from two or more of the synthetic peptides. In various embodiments, an assembled synthetic protein is assembled from two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen synthetic peptides. In various embodiments, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14. In various embodiments, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27. In various embodiments, determining interference and / or sample processing artifacts across the plurality of channels using the detected signal comprises determining ion interference. In various embodiments, the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%. In various embodiments, the one or more assembled synthetic proteins comprise one or more assembled synthetic, non-naturally occurring peptides. In various embodiments, the one or more synthetic peptides comprise one or more synthetic, non- naturally occurring peptides. In various embodiments, the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils,4IPTS / 200244493 1hamsters, rats and mice). In various embodiments, the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

[0015] Further disclosed herein is a method for determining interference and / or sample processing artifacts in mass spectrometry -based proteomics, the method comprising: generating a mixture of one or more synthetic peptides and one or more background peptides of a background proteome, wherein concentrations of each synthetic peptide is selected to account for differences in ionization and flight; performing mass spectrometry to detect signal from at least the one or more synthetic peptides, or fragments thereof, across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal. In various embodiments, the one or more synthetic peptides are between 5 and 45 amino acids in length. In various embodiments, the one or more synthetic peptides comprise between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides. In various embodiments, the mixture comprises about 20 synthetic peptides. In various embodiments, the one or more synthetic peptides are selected from a protein comprising a sequence sharing at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or 100% identity to any one of SEQ ID NOs: 1-27. In various embodiments, the mixture comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 different synthetic peptides each comprising a protein sequence selected from the group consisting of SEQ ID NOs: 1-27. In various embodiments, determining interference and / or sample processing artifacts across the plurality of channels using the detected signal comprises determining ion interference. In various embodiments, the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%.

[0016] Additionally disclosed herein is a method for detecting a processing anomaly, the method comprising: providing a mixture of one or more assembled synthetic peptides and one or more background peptides of a background proteome, the one or more assembled5IPTS / 200244493 1synthetic peptides each assembled from two or more synthetic peptides, and wherein the two or more synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; performing one or more of: digesting the mixture using protein digestion reagents; performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents; and purifying proteins or fragments thereof in the mixture; performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and detecting a processing anomaly arising from the one or more of the digestion, labeling, or purifying steps.

[0017] In various embodiments, providing the mixture comprises spiking a known concentration of the one or more assembled synthetic proteins into a sample comprising the one or more background peptides of a background proteome. In various embodiments, an assembled synthetic protein is assembled from two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen synthetic peptides. In various embodiments, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14. In various embodiments, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27. In various embodiments, an assembled synthetic protein comprises a sequence sharing at least 80% identity, at least 85% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.1% identity, at least 99.2% identity, at least 99.3% identity, at least 99.4% identity, at least 99.5% identity, at least 99.6% identity, at least 99.7% identity, at least 99.8% identity, at least 99.9% identity, or 100% identity with SEQ ID NO: 28 or SEQ ID NO: 29. In various embodiments, the two or more synthetic peptides are between 5 and 45 amino acids in length. In various embodiments, the one or more synthetic peptides comprise between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides. In various embodiments, the protein digestion reagents comprise trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin. In various embodiments, the one or more peptide labeling agents comprise tandem mass tags (TMTs).6IPTS / 200244493 1

[0018] In various embodiments, the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%. In various embodiments, the one or more assembled synthetic proteins comprise one or more assembled synthetic, non-naturally occurring peptides. In various embodiments, the two or more synthetic peptides comprise two or more synthetic, non-naturally occurring peptides. In various embodiments, the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, nonhuman primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice). In various embodiments, the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

[0019] Additionally disclosed herein is a kit comprising: one or more synthetic peptides: a background proteome comprising one or more background peptides from cells, wherein the one or more synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; protein digestion reagents; one or more peptide labeling agents; and instructions for determining interference and / or sample processing artifacts in an analytic procedure using the one or more synthetic peptides, the one or more background peptides of the background proteome, the protein digestion reagents, and the one or more peptide labeling agents. In various embodiments, the instructions further comprise instructions for; generating, or providing a generated mixture of, the one or more synthetic peptides and the one or more background peptides of the background proteome; optionally performing labeling of proteins or fragments thereof in the digested mixture using the one or more peptide labeling agents; performing a mass spectrometry analytic procedure to detect signal to noise (SN) values across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected SN values.

[0020] In various embodiments, the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%. In various7IPTS / 200244493 1embodiments, the one or more synthetic peptides are between 5 and 45 amino acids in length. In various embodiments, the one or more synthetic peptides comprise between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides. In various embodiments, one or more synthetic peptides comprise two or more synthetic, non-naturally occurring peptides. In various embodiments, the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, nonhuman primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice). In various embodiments, the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

[0021] Additionally disclosed herein is a kit for determining interference and / or sample processing artifacts in mass spectrometry -based proteomics, the method comprising: between 2 and 30 synthetic peptides between 5 and 45 amino acids in length; a background proteome comprising one or more background peptides from cells, wherein the 2 and 30 synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid: protein digestion reagents: one or more peptide labeling agents; and instructions for determining interference and / or sample processing artifacts in an analytic procedure using the one or more synthetic peptides, the one or more background peptides of the background proteome, the protein digestion reagents, and the one or more peptide labeling agents. In various embodiments, the instructions further comprise instructions for: generating, or providing a generated mixture of, the one or more synthetic peptides and the one or more background peptides of the background proteome; optionally performing labeling of proteins or fragments thereof in the digested mixture using the one or more peptide labeling agents; performing a mass spectrometry analytic procedure to detect signal to noise (SN) values across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected SN values.

[0022] In various embodiments, the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between8IPTS / 200244493 10.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%. In various embodiments, the mixture comprises about 20 synthetic peptides. In various embodiments, the one or more synthetic peptides are selected from a protein comprising a sequence sharing at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or 100% identity to any one of SEQ ID NOs: 1-27. In various embodiments, the mixture comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 different synthetic peptides each comprising a protein sequence selected from the group consisting of SEQ ID NOs: 1 -27. In various embodiments, the protein digestion reagents comprise trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin. In various embodiments, the one or more peptide labeling agents comprise tandem mass tags (TMTs). In various embodiments, an assembled synthetic protein assembled from two or more of the synthetic peptides. In various embodiments, an assembled synthetic protein is assembled from two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen synthetic peptides. In various embodiments, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14. In various embodiments, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27. In various embodiments, the 2 to 30 synthetic peptides comprise 2 to 30 synthetic, non-naturally occurring peptides. In various embodiments, the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs. Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice). In various embodiments, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

[0023] Additionally disclosed herein is a mixture for determining analytic interference, the mixture comprising: between 1 to 20 synthetic peptides each comprising a sequence selected from the group consisting of SEQ ID NOs: 1-27. In various embodiments, the mixture further comprises one or more background peptides of a background proteome. In various embodiments, the one or more background peptides of the background proteome are obtained by: lysing a plurality of cells; and extracting the one or more background peptides from the lysed plurality of cells. In various embodiments, the one or more background peptides of a9IPTS / 200244493 1background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice). In various embodiments, the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans. In various embodiments, the 1 to 20 synthetic peptides comprise 1 to 20 synthetic, non-naturally occurring peptides.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The foregoing and other objects, features and advantages will become apparent from the following description of preferred embodiments, as illustrated in the accompanying drawings. Like referenced elements identify common features in the corresponding drawings. The drawings are not necessarily to scale, with emphasis instead being placed on illustrating the principles of the present disclosure, in which:

[0025] Figure (FIG.) 1 is an example flow diagram for determining interference and / or sample processing artifacts in mass spectrometry tag-based proteomics, in accordance with an embodiment.

[0026] FIG. 2A is an example flow process for determining interference and / or sample processing artifacts in mass spectrometry tag-based proteomics, in accordance with an embodiment.

[0027] FIG. 2B is an example flow process for detecting a processing anomaly, in accordance with an embodiment.

[0028] FIG. 3 shows an example peptide selection scheme.

[0029] FIGs. 4A and 4B show signal (FIG. 4A) and signal to noise (S / N) (FIG. 4B) values across a plurality of channels for a first protein (KOI).

[0030] FIGs. 5A and 5B show signal (FIG. 5A) and signal to noise (S / N) (FIG. 4B) values across a plurality of channels for a second protein (KO2).

[0031] FIGs. 6A and 6B show signal (FIG. 6A) and signal to noise (S / N) (FIG. 6B) values across a plurality of channels for a third protein (KO3).

[0032] FIGs. 7A and 7B show signal (FIG. 7A) and signal to noise (S / N) (FIG. 7B) values across a plurality of channels for a fourth protein (KO4).10IPTS / 200244493 1

[0033] FIGs. 8A and 8B show signal (FIG. 8A) and signal to noise (S / N) (FIG. 8B) values across a plurality of channels for a fifth protein across concentrations of a dose response (DR1).

[0034] FIGs. 9A and 9B show signal (FIG. 9A) and signal to noise (S / N) (FIG. 9B) values across a plurality of channels for a sixth protein across concentrations of a dose response (DR2).

[0035] FIGs. 10A and 10B show signal (FIG. 10A) and signal to noise (S / N) (FIG. 10B) values across a plurality of channels for a seventh protein across concentrations of a dose response (DR3).

[0036] FIG. 11 shows background signal across a plurality of channels.

[0037] FIGs. 12A and 12B show no and high interferences across a plurality of channels.

[0038] FIGs. 13A and 13B show interference as a function of protein concentration (protein of SEQ ID NO: 26).

[0039] FIGs. 14A and 14B show exemplary methods for tracking processing anomalies through engineered proteins.

[0040] FIGs. 15A and 15B show different signals across a plurality of channels, which shows differences between fully tryptic peptides and missed cleavage peptides.DETAILED DESCRIPTIONDEFINITIONS

[0041] Terms used in the claims and specification are defined as set forth below unless otherwise specified.

[0042] The term "about" generally means a range that can be greater than or less than 10% of the value stated in a particular usage context. For example, "about 10" includes the range 9- 11.

[0043] The term “peptide” refers to a chain of amino acids. The term peptide includes any amino acid sequence comprising two or more consecutive amino acid residues, e.g. at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 amino acids. In various embodiments, a protein includes between 2 and 100 amino acids, between 3 and 90 amino acids, between 4 and 80 amino acids, between 5 and 70 amino acids, between 6 and 60 amino acids, between 7 and 50 amino acids, between 8 and 40 amino acids, between 9 and 30 amino acids, or between 10 and 20 amino acids. In particular embodiments, a peptide includes between 5 and 45 amino11IPTS / 200244493 1acids. In various embodiments, a peptide includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29. 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 amino acids.

[0044] The term “synthetic peptide” refers to a peptide sequence that does not overlap with sequences that are found in background peptides (also referred to herein as a “synthetic, nonbackground overlapping peptide”. In various embodiments, a first sequence does not overlap with another sequence if no X consecutive amino acid sequence is identical to another equal length portion of the second sequence. In various embodiments, X is fewer than 10 consecutive amino acids, fewer than 9 consecutive amino acids, fewer than 8 consecutive amino acids, fewer than 7 consecutive amino acids, fewer than 6 consecutive amino acids, fewer than 5 consecutive amino acids, fewer than 4 consecutive amino acids, fewer than 3 consecutive amino acids, or fewer than 2 consecutive amino acids.

[0045] The term “background proteins” refer to proteins of a background proteome obtained from a known organism. In various embodiments, the background proteome comprises at least 5000 translated proteins, at least 10000 translated proteins, at least 15000 translated proteins, or at least 20000 translated proteins, at least 30000 translated proteins, at least 40000 translated proteins, or at least 50000 translated proteins. A “translated protein” refers to a protein typically found in the known organism. In various embodiments, background proteins are proteins of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice). In particular embodiments, background proteins are proteins of a background proteome obtained from one or more humans. In such embodiments, the background proteins are human background proteins of a human background proteome.

[0046] The phrase “synthetic, non-naturally occurring peptide” refers to a synthetic peptide that is not found in nature.

[0047] The term “synthetic protein” refers to a protein sequence. Generally, a synthetic protein is larger than a synthetic peptide. For example, a synthetic protein can be assembled from two or more synthetic peptides, also referred to herein as an “assembled synthetic protein.” A synthetic protein can be 100 or more consecutive amino acid residues, e.g. at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, or at least 200 amino acids. In various embodiments, a synthetic protein is between 100 and 1000 amino acids e.g., between 110 and 900 amino12IPTS / 200244493 1acids, between 120 and 800 amino acids, between 130 and 700 amino acids, between 140 and 600 amino acids, between 150 and 500 amino acids, between 160 and 400 amino acids, between 170 and 300 amino acids, or between 180 and 200 amino acids. In various embodiments, a synthetic protein is about 110 amino acids, about 120 amino acids, about 130 amino acids, about 140 amino acids, about 150 amino acids, about 160 amino acids, about 170 amino acids, about 180 amino acids, about 190 amino acids, or about 200 amino acids.

[0048] It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise.OVERVIEW

[0049] Disclosed herein are methods for determining interference and / or sample processing artifacts in mass spectrometry-based proteomics and / or detecting a processing anomaly. In various embodiments, methods for determining interference and / or sample processing artifacts in mass spectrometry tag-based proteomics comprise providing a mixture of one or more synthetic peptides and one or more background peptides of a background proteome; performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents; performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal. In various embodiments, methods for determining interference and / or sample processing artifacts in mass spectrometry-based proteomics comprise providing a mixture comprising between 2 and 30 synthetic peptides between 5 and 45 amino acids in length and one or more background peptides of a background proteome; performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents; performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal. In various embodiments, methods for determining interference and / or sample processing artifacts in mass spectrometry tag-based proteomics comprise generating a mixture of one or more synthetic peptides and one or more background peptides of a background proteome, wherein concentrations of each synthetic peptide is selected to account for differences in ionization and flight; performing mass spectrometry to detect13IPTS / 200244493 1signal from at least the one or more synthetic non-naturally occurring proteins, or fragments thereof, across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal.

[0050] Reference is made to FIG. 1, which is an example flow diagram for determining interference and / or sample processing artifacts in mass spectrometry tag-based proteomics, in accordance with an embodiment. FIG. 1 introduces one or more synthetic peptides 105, such as synthetic non-naturally occurring peptides (e.g., synthetic peptides that are not found in nature). FIG. 1 depicts an embodiment including three synthetic peptides 105; however, in some embodiments, there may be fewer or more synthetic peptides 105. In various embodiments, synthetic peptides 105 may include between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides. In particular embodiments, synthetic peptides 105 include 20 peptides. In various embodiments, a synthetic peptide 105 includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 amino acids. In various embodiments, a synthetic peptide 105 includes between 2 and 100 amino acids, between 3 and 90 amino acids, between 4 and 80 amino acids, between 5 and 70 amino acids, between 6 and 60 amino acids, between 7 and 50 amino acids, between 8 and 40 amino acids, between 9 and 30 amino acids, or between 10 and 20 amino acids. In particular embodiments, a synthetic peptide 105 includes between 5 and 45 amino acids.

[0051] FIG. 1 further introduces one or more cells 110. In various embodiments, the cells 110 are obtained from a human donor. For example, a sample can be obtained from a human donor and one or more cells 110 can be isolated from the sample. As used herein, the term “sample” can include a single cell or multiple cells or fragments of cells or an aliquot of body fluid, such as a blood sample, taken from a subject, by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, or intervention or other means known in the art.

[0052] In various embodiments, the cells 110 are from a maintained cell line. In various embodiments, the cells 110 can be obtained from any of a unicellular organism (e.g., bacteria or yeast) or a multicellular organism (e.g., a mammalian organism such as humans, non-14IPTS / 200244493 1human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice). In particular embodiments, the cells 110 are yeast cells. In particular embodiments, the cells 110 are Expi293 cells, though any human cell line can be appropriately used. In various embodiments, the cells 110 can be any of a single cell, population of cells, or multiple populations of cells, and can further vary in regards to the type of cells (single cell type, mixture of cell types), cell culture (e.g., in vivo, in vitro 2D culture, or in vitro in vitro 3D culture). In various embodiments, the cells 110 are obtained from a healthy individual, such that the resulting human proteins of the background proteome represent background proteins from a healthy individual (and therefore can be appropriate for use in a standard for mass spectrometry).

[0053] As shown in FIG. 1, the cells 110 are exposed to lysis conditions 112 which cause the cells 110 to undergo lysis, resulting in lysed cells 120. Lysis conditions 112 include any suitable condition that causes the cells 1 10 to lyse. Lysis conditions 1 12 include cell lysis reagents, examples of which include detergents (e.g., sodium dodecyl sulphate, Tween, Triton-X, NP-40) and enzymes (e.g.. lysozyme, lysostaphin, maltase, cellulase, mutanolysin, glycanase, protease, mannanase, proteinase K). Lysis conditions 112 can additionally or alternatively include cellular disruption means including electroporation, thermal disruption, acoustic disruption, mechanical disruption, sonication, liquid homogenization, or through use of osmotic pressure, e.g., using a hypotonic lysis buffer. In some embodiments, a cell may be lysed via formation of a gel matrix (e.g., within the cell, such as by use of a cross-linking agent).

[0054] As shown in FIG. 1, the lysed cells 120 undergo protein labeling and / or extraction 122 to obtain the background peptides 130. In various embodiments, protein labeling refers to labeling specific amino acids of proteins. For example, protein labeling can include labeling thiol groups of cysteine. An example agent for labeling specific amino acids (e.g., labeling thiol groups of cysteines) is desthiobiotin iodoacetamide (DBIA). Various protein extraction methods can be used for performing protein extraction, examples of which include differential centrifugation, protein precipitation (e.g., immunoprecipitation or ammonium sulfate precipitation), protein chromatography (e.g., hydrophobic interaction chromatography, size exclusion chromatography, ion exchange chromatography), gel electrophoresis, proteome fractionation, or affinity purification.15IPTS / 200244493 1

[0055] In various embodiments, although not explicitly shown in FIG. 1, methods can involve spiking one or more assembled synthetic proteins into a sample obtained after any of step 112 and / or step 122. For example, methods can involve spiking one or more assembled synthetic proteins into one or more of (i) a sample of the lysed plurality of cells (e.g., lysed cells 120), (ii) a sample of labeled background peptides (e.g., thiol-labeled proteins after protein labeling at step 122), and (iii) a sample of the extracted one or more background peptides (e.g., proteins after extraction at step 122). As discussed in further detailed herein, spiking assembled synthetic proteins can be useful for determining a processing anomaly e.g., processing anomaly associated with any of the steps (e.g., including protein labeling or extraction at step 122).

[0056] As shown in FIG. 1, the background proteins 130 undergo digestion 138, thereby generating digested background peptides 140. Digestion 138 can involve providing one or more protein digestion reagents to the mixture 135. Example protein digestion reagents include any of trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin.

[0057] In various embodiments, protein digestion reagents are provided at a particular ratio relative to the background proteins 130. In various embodiments, protein digestion reagents are provided at a final protein digestion reagent to background protein ratio between 1:10 to 1:500 (w / w). In various embodiments, protein digestion reagents are provided at a final protein digestion reagent to background protein ratio between 1 :15 to 1:250 (w / w), between 1 :20 to 1 :100 (w / w), between 1 :30 to 1 :90 (w / w), between 1 :40 to 1 :75 (w / w), or between 1:50 to 1:60 (w / w). The mixture 135 can be exposed to conditions to facilitate the digestion, an example of which includes elevated temperature. In various embodiments, the mixture 135 is exposed to a temperature between 35°C and 50°C. In various embodiments, the mixture 135 is exposed to a temperature between 35°C and 45°C, between 35°C and 40°C, or between 36°C and 38°C. In particular embodiments, the mixture is exposed to a temperature of about 37°C.

[0058] The synthetic peptides 105 and the digested background peptides 140 are combined to generate a mixture 135. Here, the mixture contains one or more synthetic peptides and one or more background peptides of the background proteome. In various embodiments, the mixture comprises between 2 and 30 synthetic peptides. In various embodiments, the mixture comprises between 4 and 29 synthetic peptides, between 6 and 28 synthetic peptides, between 8 and 27 synthetic peptides, between 10 and 26 synthetic peptides, between 12 and 25 synthetic peptides, between 14 and 24 synthetic peptides, between 16 and 23 synthetic16IPTS / 200244493 1peptides, between 18 and 22 synthetic peptides, or between 19 and 21 synthetic peptides. In various embodiments, each synthetic peptide is between 5 and 45 amino acids in length.

[0059] In various embodiments, the mixture 135 includes a target concentration of synthetic peptides relative to the background peptides. In various embodiments, the target concentration of a synthetic peptide is between 0.0001 % and 0. 1 %. In various embodiments, the target concentration of a synthetic peptide when mixed with the background peptides is between 0.0002% and 0.05%, between 0.0003% and 0.04%, between 0.0004% and 0.03%, between 0.00045% and 0.02%, or between 0.0005% and 0.01%. In various embodiments, the mixture 135 comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%. In various embodiments, the target concentration of a synthetic peptide when mixed with the background peptides is between 0.0005% and 0.01%. In various embodiments, the target concentration of a synthetic peptide when mixed with the human peptides is 0.0001%, 0.0002%, 0.0003%, 0.0004%, 0.0005%, 0.0006%, 0.0007%, 0.0008%, 0.0009%, 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%. 0.007%, 0.008%, 0.009%, 0.01%, 0.011%, 0.012%, 0.013%, 0.014%, 0.015%, 0.016%, 0.017%, 0.018%, 0.019%, or 0.020%.

[0060] In various embodiments, the mixture 135 comprises between 1 to 20 synthetic peptides each comprising a sequence selected from the group consisting of SEQ ID NOs: 1 - 27. In various embodiments, the mixture 135 comprises between 1 to 20 synthetic peptides each comprising a sequence sharing at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94 identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or 100% identity to any one of SEQ ID NOs: 1-27.

[0061] As shown in FIG. 1, the mixture 135 undergoes peptide labeling 142. In various embodiments, peptide labeling 142 involves performing isobaric labeling. For example, isobaric labeling involves providing the mixture 135 with tags. Various suitable isobaric mass labels are known in the art such as Tandem Mass Tags (Thompson et al., 2003, Anal. Chem. 75(8): 1895-1904 (incorporated herein by reference in its entirety) disclosed in WO 01 / 68664 (incorporated herein by reference in its entirety) and WO 03 / 025576 (incorporated herein by reference in its entirety), iPROT tags disclosed in U.S. Pat. No. 6,824,981 (incorporated herein by reference in its entirety) and iTRAQ tags (Pappin et al.,17IPTS / 200244493 12004, Methods in Clinical Proteomics Manuscript M400129-MCP200 (incorporated herein by reference in its entirety)). The term “isobaric” means that the mass labels have substantially the same aggregate mass as determined by mass spectrometry. Typically, the average molecular masses of the isobaric mass labels will fall within a range of ±0.5 Da of each other.

[0062] Isobaric labeling using either tandem mass tags (TMT) or isobaric tags for relative and absolute quantitation (iTRAQ) uses mass spectrometry to quantitate and compare proteins for quantitation of peptides. TMT and iTRAQ have the general structure M - F - N - R where M = mass reporter region, F = cleavable linker region, N = mass normalization region, and R = protein reactive group. Isotopes substituted at various positions in M and N cause each tag to have a different molecular mass in the M region with a corresponding mass change in the N region, so that the set of tags have the same overall molecular weight. When the TMT undergo a second or third fragmentation (such as in tandem mass spectrometry MS / MS, or triple mass spectroscopy MS / MS / MS) the fragmented tags are distinguishable, with backbone fragmentation yielding sequence and tag fragmentation yielding mass reporter ions for quantitating the peptides. iTRAQ and TMT covalently label amine groups in protein digests and a cysteine reactive TMT labels thiols of cysteines, resulting in individual digests with unique mass tags. The labeled digests are then pooled and fragmented into peptide backbone and reporter ions. The peptide backbone ions are used to identify the protein from which they came. The reporter ions are used to quantify this protein in each of the combined samples.

[0063] Returning to FIG. 1, the protein labeling 142 process generates tagged peptides / fragments 145. Mass spectrometry 150 is performed to analyze the tagged peptides / fragments 145. Here, the tagged peptides / fragments 145 may serve as an internal standard (given that concentrations of proteins / fragments are known in the mixture 135). Further details for performing mass spectrometry (MS) are described herein.

[0064] Generally, the tagged peptides / fragments 145 serving as the internal standards are characterized according to their mass-to-charge ratio (m / z). In various embodiments, the tagged peptides / fragments 145 serving as the internal standards are characterized according to their retention time on a chromatographic column (e.g., such as an HPLC column). In various embodiments, the peptide internal standard is analyzed by MSI, which is a first level of mass spectrometry that surveys all intact proteins and / or fragments in the sample. The tagged peptides / fragments 145 can be further fragmented with multistage MS (MSn).18IPTS / 200244493 1Fragmentation can be achieved by inducing ion / molecule collisions by a process known as collision-induced dissociation (CID) (also known as collision-activated dissociation (CAD)). Collision-induced dissociation is accomplished by selecting a peptide ion of interest with a mass analyzer and introducing that ion into a collision cell. The selected ion then collides with a collision gas (typically argon or helium) resulting in fragmentation. Generally, any method that is capable of fragmenting a peptide is encompassed within the scope of the disclosed methods. In addition to CID, other fragmentation methods include, but are not limited to, high energy collisional dissociation (HCD), surface induced dissociation (SID), blackbody infrared radiative dissociation (BIRD), electron capture dissociation (ECD), postsource decay (PSD), LID, and the like. The fragments are then analyzed to obtain a fragment ion spectrum.

[0065] In some embodiments, the tagged peptides / fragments 145 are analyzed by at least two stages of mass spectrometry to determine the fragmentation pattern of the peptides and to identify a peptide fragmentation signature for peptide verification. A peptide signature is obtained in which peptide fragments have significant differences in m / z ratios to enable peaks corresponding to each fragment to be well separated. Signatures are desirably unique, i.e., diagnostic of a peptide being identified and comprising minimal if any overlap with fragmentation patterns of peptides with different amino acid sequences. If a suitable fragment signature is not obtained at the first stage, additional stages of mass spectrometry are performed until a unique signature is obtained. This fragmentation signature ensures that peaks of the same exact mass are not simply rearrangements of the same amino acids in a different order.

[0066] The output of the mass spectrometry 150 may be a spectrum (or spectra), such as a fragment ion spectrum reflecting quantities of resulting ions (resulting ions from the tagged peptides / fragments 145). In various embodiments, the spectrum can be expressed as a plurality of channels where each channel corresponds to a distinguishable portion of a peptide labeling agent (e.g., TMT). As shown below in the Examples, the spectrum includes different signal (or signal to noise) values across the plurality of channels. For example, example channels can correspond to any of 126, 127n, 127c, 128n, 128c, 129n, 129c, 130n, 130c, 131n, 131c, 132n, 132c, 133n, 133c, 134n, 134c, and 135.

[0067] The output of the mass spectrometry 150 is used to determine interference e.g., ion interference. In various embodiments, interference can stem from a variety of sources, including MS instrument-specific interferences or non-proteinaceous interference (e.g.,19IPTS / 200244493 1contaminants, adducts, solvents, or polymeric interference). Thus, determining interference and / or sample processing artifacts can be useful for ensuring accuracy in mass spectrometry tag-based proteomics.

[0068] In various embodiments, determining interference and / or sample processing artifacts involves comparing the signal (or signal to noise) values of at least two spectra. For example, a first spectrum may be a knockout (KO) spectrum may include signal in a first set of channels and lack of signal in a second set of channels. As the mass of the TMTs and synthetic, non-occurring proteins are known and fragments thereof are not expected to be detected in the second set of channels, this theoretical result is expected. A “knockout” (KO) can refer to the lack of a protein or fragment thereof and therefore signals in the one or more other channels are not expected to be present given the KO. A second spectrum may be a different spectrum that includes one or more synthetic, non-occurring proteins that are known to be detectable in the second set of channels. Thus, the second spectrum may differ from the first spectrum in at least the signal in the second set of channels. The signal values in the second set of channels (referred to as the “knockout” channels) in the first spectrum and the second spectrum can be useful for determining the interference. For example, if there is signal in the second set of channels in the first spectrum (where signal is not expected to be detected), such signal can be attributed at least to interference e.g., ion interference.

[0069] Although the foregoing description is in relation to absolute signal values (e.g., channels include or do not include signal), the same principle can be applied for determining interference and / or sample processing artifacts when relative signal values are present in the channels. For example, in some scenarios, the the synthetic, non-occurring proteins are provided at different concentrations (e.g., of a dose response (DR)). Thus, the theoretical signal in certain channels can be expected to increase in accordance with the increasing concentrations. For example, if the concentrations are increasing linearly, then the signal can similarly be expected to increase linearly. In some scenarios, the spectrum in the channels does not change as expected. For example, the spectrum may be less linear or non-linear. Thus, the detected relative signal values across the channels can be attributed at least to interference.

[0070] In various embodiments, determining interference involves calculating a metric, such as a interference free index (IFI). The IFI is determined based on intensities of reporter ions (corresponding to each channel of the spectrum) and intensities of the non- KO channels.20IPTS / 200244493 1Thus, an IFI value closer to 1 reflects little to no interference whereas an IFI value closer to 0 indicates higher interference. In various embodiments, IFI can be calculated as:where S / N (KO) refers to the average signal to noise (S / N) values of the knockout channels and S / N (notKO) refers to the average SN values of the non-sKO channels. Further details of an IFI metric is described in Navarrete-Perea et al., “Synthetic Knockout Protein Standard for Evaluating Interference in Tandem Mass Tag-Based Proteomics,” Analytical Chemistry (2024), 96(17)), which is incorporated by reference in its entirety.

[0071] Returning to FIG. 1, the determined interference 155 can useful such that the subsequent mass spectrometry measurements can account for the interference 155. For example, for subsequent multiplexing proteomics experiments (e.g., which uses isobaric labeling), the resulting mass spectrometry measurements can correct for the determined interference 155, thereby ensuring that the measurements for the multiplexed proteomics experiments are more accurate.METHODS FOR DETERMINING INTERFERENCE IN MASS SPECTROMETRY TAG- BASED PROTEOMICS

[0072] FIG. 2A is an example flow process for determining interference in mass spectrometry tag-based proteomics, in accordance with an embodiment.

[0073] Step 205 involves providing one or more background proteins of a background proteome.

[0074] Step 210 involves digesting the one or more background proteins of the background proteome using protein digestion reagents, examples of which include trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin. In particular embodiments, step 210 involves digesting the mixture using trypsin.

[0075] Step 215 involves combining the background proteome with one or more synthetic peptides to generate a mixture.

[0076] Step 220 is an optional step, as indicated by the dotted lines shown in FIG. 2. Step 220 involves performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents. In particular embodiments, step 220 involves performing isobaric labeling e.g., using tandem mass tags (TMTs).21IPTS / 200244493 1

[0077] Step 225 involves performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a synthetic peptide or a fragment thereof.

[0078] Step 230 involves determining interference across the plurality of channels using the detected signal.METHODS FOR DETECTING A PROCESSING ANOMALY IN MASS SPECTROMETRY PROTEOMICS

[0079] As disclosed herein, methods may involve detecting a processing anomaly. In some embodiments, the method for detecting a processing anomaly involves the use of one or more assembled synthetic proteins. In various embodiments, an assembled synthetic protein is assembled from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 synthetic peptides. In various embodiments an assembled synthetic protein is assembled from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 synthetic peptides.

[0080] As an example, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14. As another example, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27.

[0081] In various embodiments, an assembled synthetic protein shares at least 80% identity, at least 85% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.1% identity, at least 99.2% identity, at least 99.3% identity, at least 99.4% identity, at least 99.5% identity, at least 99.6% identity, at least 99.7% identity, at least 99.8% identity, at least 99.9% identity, or 100% identity with SEQ ID NO: 28 or SEQ ID NO: 29.

[0082] Reference is made to FIG. 2B, which is an example flow process for detecting a processing anomaly, in accordance with an embodiment.

[0083] Step 255 involves providing a mixture of one or more assembled synthetic proteins and one or more background peptides of a background proteome, the one or more assembled synthetic proteins each assembled from two or more synthetic peptides. In various embodiments, providing the mixture comprises spiking a known concentration of the one orIPTS / 200244493 1more assembled synthetic proteins into a sample comprising the one or more background peptides of a background proteome.

[0084] As shown in FIG. 2B, one or more of steps 260, 265, and 270 can be performed. In some embodiments, only one of steps 260, 265, and 270 is performed. In some embodiments, two of steps 260, 265, and 270 are performed. In some embodiments, each of steps 260, 265, and 270 are performed.

[0085] Step 260 involves digesting the mixture using protein digestion reagents, examples of which include trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin.

[0086] Step 265 involves performing labeling (e.g., isobaric labeling) of proteins or fragments thereof in the mixture using one or more peptide labeling agents (e.g., TMTs).

[0087] Step 270 involves purifying proteins or fragments thereof in the mixture.

[0088] Step 275 involves performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent.

[0089] Step 280 involves detecting a processing anomaly arising from the digestion, labeling, and / or purifying step. In various embodiments, a processing anomaly refers to an inefficiency associated with one or more processing steps, such as one of the digestion, labeling, and / or purifying steps. For example, a processing anomaly can be a digestion efficiency of proteins (e.g., where a portion of proteins are successfully cleaved and a second portion of proteins are not cleaved). As another example, a processing anomaly can be a labeling efficiency of proteins (e.g., where a portion of proteins are successfully labeled and a second portion of proteins are not labeled).

[0090] In various embodiments, one or more assembled synthetic proteins need not be provided at step 255 into the mixture. Instead, the one or more assembled synthetic proteins can be spiked into a sample following one of steps 260, 265, or 270. For example, one or more assembled synthetic proteins can be spiked into one or more of: (i) a sample of the digested mixture from step 260, (ii) a sample of the labeled proteins or fragments thereof from step 265, or (iii) a sample of purified proteins or fragments thereof from step 270. By analyzing, through mass spectrometry at step 275, individual processing anomalies attributable to each of step 260, 265, or step 270 can be determined.EXEMPLARY SYNTHETIC PEPTIDES AND ASSEMBLED SYNTHETIC PROTEINS23IPTS / 200244493 1

[0091] Disclosed herein are synthetic peptides useful as internal standards for performing mass spectrometry. As disclosed herein, such synthetic peptides can be used as internal standards for determining interference in mass spectrometry tag-based proteomics and / or for detecting a processing anomaly.

[0092] In various embodiments, synthetic peptides disclosed herein are non-naturally occurring, hereafter referred to as synthetic, non-naturally occurring proteins. Synthetic peptides can be preferable for use as internal standards, as they can be designed to avoid overlapping with human peptide sequences in a sample of interest. Furthermore, non- naturally occurring proteins can be selected and designed to ensure that they ionize and can be detectable (e.g., fragments of proteins can be detectable following digestion). In various embodiments, a synthetic peptide differs from the background peptides of the background proteome by at least 1 amino acid. In various embodiments, a synthetic peptide differs from the background peptides of the background proteome by at least 2 amino acids, by at least 3 amino acids, by at least 4 amino acids, by at least 5 amino acids, by at least 6 amino acids, by at least 7 amino acids, by at least 8 amino acids, by at least 9 amino acids, or by at least 10 amino acids.

[0093] In various embodiments, a synthetic peptide includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 amino acids. In various embodiments, a synthetic peptide includes between 2 and 100 amino acids, between 3 and 90 amino acids, between 4 and 80 amino acids, between 5 and 70 amino acids, between 6 and 60 amino acids, between 7 and 50 amino acids, between 8 and 40 amino acids, between 9 and 30 amino acids, or between 10 and 20 amino acids. In particular embodiments, a synthetic peptide includes between 5 and 45 amino acids.

[0094] In various embodiments, one or more synthetic peptides can be used as internal standards for performing mass spectrometry. In various embodiments, the one or more synthetic peptides include between 1 and 1000 peptides. In various embodiments, the one or more synthetic peptides include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20,21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35. 36, 37, 38, 39, 40, 41, 42, 43, 44, 45,46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70,71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95,96, 97, 98, 99, or 100 peptides. In various embodiments, the one or more synthetic peptides include between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides,24IPTS / 200244493 1between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides.

[0095] In various embodiments, a synthetic peptide comprises a sequence sharing at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94 identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or 100% identity to any one of SEQ ID NOs: 1-27.

[0096] In various embodiments, a synthetic protein can be assembled from two or more synthetic peptides disclosed herein. A synthetic peptide that is assembled from two or more synthetic peptides is referred to herein as an assembled synthetic protein. An assembled synthetic protein is useful for detecting a processing anomaly. For example, given the longer length of the assembled synthetic protein, the processing anomaly arising from a processing step, such as a protein digestion step, can be detected based on whether the assembled synthetic protein is successfully digested, cleaved, and detected.

[0097] In various embodiments, an assembled synthetic protein is assembled from two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, twenty one, twenty two, twenty three, twenty four, twenty five, twenty six, twenty seven, twenty eight, twenty nine, or thirty synthetic peptides. In various embodiments, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14. In various embodiments, an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27.

[0098] In various embodiments, an assembled synthetic protein comprises a sequence sharing at least 80% identity, at least 85% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.1% identity, at least 99.2% identity, at least 99.3% identity, at least 99.4% identity, at least 99.5% identity, at least 99.6% identity, at least 99.7% identity, at least 99.8% identity, at least 99.9% identity, or 100% identity with SEQ ID NO: 28 or SEQ ID NO: 29.MASS SPECTROMETRY

[0099] As disclosed herein, mass spectrometry can be performed on labeled proteins or fragments thereof (e.g., labeled through isobaric labeling). Mass spectrometry generally25IPTS / 200244493 1works by ionizing chemical compounds to generate charged molecules or molecule fragments and measuring their mass-to-charge ratios. A typical mass spectrometry procedure may include loading a sample onto the mass spectrometry instrument followed by vaporization, ionization of the sample components by any one of a variety of methods (e.g., impacting with an electron beam), resulting in charged particles (ions), separation of ions according to their mass-to-charge ratio in an analyzer by electromagnetic fields, detection of ions (e.g., by a quantitative method), and processing of the ion signal into mass spectra. Mass spectrometry methods are well known in the art (see, e.g., Burlingame et al. Anal. Chem. 70:647R-716R (1998)), and include, for example, quadrupole mass spectrometry, ion trap mass spectrometry, time-of-flight mass spectrometry, gas chromatography mass spectrometry, and tandem mass spectrometry. The basic processes associated with a mass spectrometry method are the generation of gas-phase ions, and the measurement of their mass. The movement of gas-phase ions can be precisely controlled using electromagnetic fields generated in the mass spectrometer. The movement of ions in these electromagnetic fields is proportional to the m / z (mass to charge ratio) of the ion and this forms the basis of measuring the m / z and therefore the mass of a sample. The movement of ions in these electromagnetic fields allows for the containment and focusing of the ions which accounts for the high sensitivity of mass spectrometry. During the course of m / z measurement, ions are transmitted with high efficiency to particle detectors that record the arrival of these ions. The quantity of ions at each m / z is demonstrated by peaks on a graph where the x axis is m / z and the y axis is relative abundance. Different mass spectrometers have different levels of resolution, that is, the ability to resolve peaks between ions closely related in mass. The resolution is defined as R=m / delta m, where m is the ion mass and delta m is the difference in mass between two peaks in a mass spectrum. For example, a mass spectrometer with a resolution of 1000 can resolve an ion with a m / z of 100.0 from an ion with a m / z of 100.1. Certain mass spectrometry methods can utilize various combinations of ion sources and mass analyzers which allows for flexibility in designing customized detection protocols. In some embodiments, mass spectrometers can be programmed to transmit all ions from the ion source into the mass spectrometer either sequentially or at the same time. In some embodiments, a mass spectrometer can be programmed to select ions of a particular mass for transmission into the mass spectrometer while blocking other ions.

[0100] Various mass spectrometers can be used for the methods disclosed herein. In various embodiments, a mass spectrometer includes: a sample inlet, an ion source, a mass analyzer, a detector, a vacuum system, and instrument-control system, and a data26IPTS / 200244493 1system. Difference in the sample inlet, ion source, and mass analyzer generally define the type of instrument and its capabilities. For example, an inlet can be a capillary-column liquid chromatography source or can be a direct probe or stage such as used in matrix-assisted laser desorption. Common ion sources are, for example, electrospray, including nanospray and microspray or matrix-assisted laser desorption. Mass analyzers include, for example, a quadrupole mass filter, ion trap mass analyzer and time-of-flight mass analyzer.

[0101] The ion formation process is a starting point for mass spectrum analysis. Several ionization methods are available and the choice of ionization method depends on the sample used for analysis. For example, for the analysis of polypeptides a relatively gentle ionization procedure such as electrospray ionization (ESI) can be desirable. For ESI, a solution containing the sample is passed through a fine needle at high potential which creates a strong electrical field resulting in a fine spray of highly charged droplets that is directed into the mass spectrometer. Other ionization procedures include, for example, fast-atom bombardment (FAB) which uses a high-energy beam of neutral atoms to strike a solid sample causing desorption and ionization. Matrix-assisted laser desorption ionization (MALDI) is a method in which a laser pulse is used to strike a sample that has been crystallized in an UV- absorbing compound matrix (e.g., 2,5-dihydroxybenzoic acid, alpha-cyano-4- hydroxycinammic acid, 3 -hydroxypicolinic acid (3 -HP A), di-ammoniumcitrate (DAC) and combinations thereof). Other ionization procedures known in the art include, for example, plasma and glow discharge, plasma desorption ionization, resonance ionization, and secondary ionization.

[0102] Ion mobility mass (IM) spectrometry is a gas-phase separation method. IM separates gas-phase ions based on their collision cross-section and can be coupled with time- of-flight (TOF) mass spectrometry. IM-MS is discussed in more detail by Verbeck et al. in the Journal of Biomole cular Techniques (Vol 13, Issue 2, 56-61).

[0103] Quadrupole mass spectrometry utilizes a quadrupole mass filter or analyzer. This type of mass analyzer is composed of four rods arranged as two sets of two electrically connected rods. A combination of rf and de voltages are applied to each pair of rods which produces fields that cause an oscillating movement of the ions as they move from the beginning of the mass filter to the end. The result of these fields is the production of a high- pass mass filter in one pair of rods and a low-pass filter in the other pair of rods. Overlap between the high-pass and low-pass filter leaves a defined m / z that can pass both filters and traverse the length of the quadrupole. This m / z is selected and remains stable in the quadrupole mass filter while all other m / z have unstable trajectories and do not remain in27IPTS / 200244493 1the mass filter. A mass spectrum results by ramping the applied fields such that an increasing m / z is selected to pass through the mass filter and reach the detector. In addition, quadrupoles can also be set up to contain and transmit ions of all m / z by applying a rf-only field. This allows quadrupoles to function as a lens or focusing system in regions of the mass spectrometer where ion transmission is needed without mass filtering.

[0104] A quadrupole mass analyzer, as well as the other mass analyzers described herein, can be programmed to analyze a defined m / z or mass range. Since the desired mass range of nucleic acid fragment is known, in some instances, a mass spectrometer can be programmed to transmit ions of the projected correct mass range while excluding ions of a higher or lower mass range. The ability to select a mass range can decrease the background noise in the assay and thus increase the signal-to-noise ratio. Thus, in some instances, a mass spectrometer can accomplish a separation step as well as detection and identification of certain mass-distinguishable nucleic acid fragments.

[0105] Ion trap mass spectrometry utilizes an ion trap mass analyzer. Typically, fields are applied such that ions of all m / z are initially trapped and oscillate in the mass analyzer. Ions enter the ion trap from the ion source through a focusing device such as an octapole lens system. Ion trapping takes place in the trapping region before excitation and ejection through an electrode to the detector. Mass analysis can be accomplished by sequentially applying voltages that increase the amplitude of the oscillations in a way that ejects ions of increasing m / z out of the trap and into the detector. In contrast to quadrupole mass spectrometry, all ions are retained in the fields of the mass analyzer except those with the selected m / z. Control of the number of ions can be accomplished by varying the time over which ions are injected into the trap.

[0106] Time-of-flight mass spectrometry utilizes a time-of-flight mass analyzer. Typically, an ion is first given a fixed amount of kinetic energy by acceleration in an electric field (generated by high voltage). Following acceleration, the ion enters a field-free or “drift” region where it travels at a velocity that is inversely proportional to its m / z. Therefore, ions with low m / z travel more rapidly than ions with high m / z. The time required for ions to travel the length of the field-free region is measured and used to calculate the m / z of the ion.

[0107] Gas chromatography mass spectrometry often can analyze a target in realtime. The gas chromatography (GC) portion of the system separates the chemical mixture into pulses of analyte and the mass spectrometer (MS) identifies and quantifies the analyte.

[0108] Tandem mass spectrometry can utilize combinations of the mass analyzers described above. Tandem mass spectrometers can use a first mass analyzer to separate ions28IPTS / 200244493 1according to their m / z in order to isolate an ion of interest for further analysis. The isolated ion of interest is then broken into fragment ions (called collisionally activated dissociation or collisionally induced dissociation) and the fragment ions are analyzed by the second mass analyzer. These types of tandem mass spectrometer systems are called tandem in space systems because the two mass analyzers are separated in space, usually by a collision cell. Tandem mass spectrometer systems also include tandem in time systems where one mass analyzer is used, however the mass analyzer is used sequentially to isolate an ion, induce fragmentation, and then perform mass analysis.

[0109] Mass spectrometers in the tandem in space category have more than one mass analyzer. For example, a tandem quadrupole mass spectrometer system can have a first quadrupole mass filter, followed by a collision cell, followed by a second quadrupole mass filter and then the detector. Another arrangement is to use a quadrupole mass filter for the first mass analyzer and a time-of-flight mass analyzer for the second mass analyzer with a collision cell separating the two mass analyzers.Other tandem systems are known in the art including reflectron-time-of-flight, tandem sector and sector-quadrupole mass spectrometry.

[0110] Mass spectrometers in the tandem in time category have one mass analyzer that performs different functions at different times. For example, an ion trap mass spectrometer can be used to trap ions of all m / z. A series of rf scan functions are applied which ejects ions of all m / z from the trap except the m / z of ions of interest. After the m / z of interest has been isolated, an rf pulse is applied to produce collisions with gas molecules in the trap to induce fragmentation of the ions. Then the m / z values of the fragmented ions are measured by the mass analyzer. Ion cyclotron resonance instruments, also known as Fourier transform mass spectrometers, are an example of tandem-in-time systems.

[0111] Several types of tandem mass spectrometry experiments can be performed by controlling the ions that are selected in each stage of the experiment. The different types of experiments utilize different modes of operation, sometimes called “scans,” of the mass analyzers. In a first example, called a mass spectrum scan, the first mass analyzer and the collision cell transmit all ions for mass analysis into the second mass analyzer. In a second example, called a product ion scan, the ions of interest are mass-selected in the first mass analyzer and then fragmented in the collision cell. The ions formed are then mass analyzed by scanning the second mass analyzer. In a third example, called a precursor ion scan, the first mass analyzer is scanned to sequentially transmit29IPTS / 200244493 1the mass analyzed ions into the collision cell for fragmentation. The second mass analyzer mass-selects the product ion of interest for transmission to the detector. Therefore, the detector signal is the result of all precursor ions that can be fragmented into a common product ion.

[0112] For quantification, internal standards, such as internal standards constructed from one or more synthetic peptides, may be used which can provide a signal in relation to the amount of proteins, for example, that is present or is introduced. Internal standards allow conversion of relative mass signals into absolute quantities. This can be accomplished by addition of a known quantity of a mass tag or mass label to each sample before detection of the target proteins.

[0113] A separation step sometimes can be implemented to remove salts, enzymes, or other buffer components from a protein sample. Several methods well known in the art, such as chromatography, gel electrophoresis, or precipitation, can be used to clean up the sample. For example, size exclusion chromatography or affinity chromatography can be used to remove salt from a sample. The choice of separation method can depend on the amount of a sample. For example, when small amounts of sample are available or a miniaturized apparatus is used, a micro-affinity chromatography separation step can be used. In addition, whether a separation step is desired, and the choice of separation method, can depend on the detection method used. Salts can absorb energy from the laser in matrix-assisted laser desorption / ionization and result in lower ionization efficiency. Thus, the efficiency of matrix- assisted laser desorption / ionization and electrospray ionization sometimes can be improved by removing salts from a sample.EXEMPLARY KITS

[0114] Disclosed herein is a kit comprising one or more synthetic peptides disclosed herein. In various embodiments, kits are useful for performing multiplex proteomics using mass spectrometry. In various embodiments, kits are useful for determining interference in mass spectrometry tag-based proteomics using the one or more synthetic peptides. In various embodiments, kits are useful for detecting a processing anomaly using one or more synthetic peptides.

[0115] In various embodiments, the kit comprises between 1 and 1000 synthetic peptides. In various embodiments, the kit comprises between 2 and 100 synthetic peptides, between 4 and 90 synthetic peptides, between 6 and 80 synthetic peptides, between 8 and 7030IPTS / 200244493 1synthetic peptides, between 10 and 60 synthetic peptides, between 12 and 50 synthetic peptides, between 14 and 40 synthetic peptides, between 16 and 30 synthetic peptides, or between 18 and 22 synthetic peptides. In various embodiments, the kit comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 different synthetic peptides. In various embodiments, the different synthetic peptides each comprise a protein sequence selected from the group consisting of SEQ ID NOs: 1-27.

[0116] In various embodiments, the kit comprises one or more synthetic peptides, each synthetic peptide stored separately from other synthetic peptides. In various embodiments, a synthetic peptide is included in the kit in a stable, lyophilized form. In various embodiments, a synthetic peptide is included in the kit in solubilized form (e.g., within a solution).

[0117] In various embodiments, a kit comprises an assembled synthetic protein assembled from two or more of the synthetic peptides. In various embodiments, a kit comprises an assembled synthetic protein that is assembled from two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen synthetic peptides. In various embodiments, the kit comprises an assembled synthetic protein that is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14. In various embodiments, the kit comprises an assembled synthetic protein that is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27.

[0118] In various embodiments, a kit comprises a background proteome comprising one or more human peptides from cells. In various embodiments, one or more human peptides from cells are included in the kit in a stable, lyophilized form. In various embodiments, one or more human peptides from cells are included in the kit in solubilized form (e.g., within a solution).

[0119] In various embodiments, a kit comprises cells containing one or more human peptides. Thus, human peptides may be obtained from the cells (e.g., via cell lysis and / or protein extraction). In various embodiments, a kit includes cells and further includes cell lysis reagents, examples of which include detergents (e.g., sodium dodecyl sulphate, Tween, Triton-X, NP-40) and enzymes (e.g., lysozyme, lysostaphin, maltase, cellulase, mutanolysin, glycanase, protease, mannanase, proteinase K).

[0120] In various embodiments, a kit comprises one or more synthetic peptides at a quantity of 5 pg, 6 pg, 7 pg, 8 pg, 9 pg, 10 pg, 11 pg, 12 pg, 13 pg, 14 pg, 15 pg, 16 pg, 1731IPTS / 200244493 1pg, 18 pg, 19 pg, 20 pg, 21 pg, 22 pg, 23 pg, 24 pg, 25 pg, 26 pg, 27 pg, 28 pg, 29 pg, 30 pg, 35 pg, 40 pg, 45 pg, 50 pg, 55 pg, 60 pg, 65 pg, 70 pg, 75 pg, 80 pg, 85 pg, 90 pg, 95 pg, 100 pg, 150 pg, 200 pg, 300 pg, 400 pg, 500 pg, 600 pg, 700 pg, 800 pg, 900 pg, or 1000 pg. In various embodiments, a kit comprises one or more synthetic peptides at a quantity between 1 pg and 1 mg. In various embodiments, a kit comprises one or more synthetic peptides at a quantity between 5 pg and 0.9 mg, between 6 pg and 0.8 mg, between 7 pg and 0.7 mg, between 8 pg and 0.6 mg, between 9 pg and 0.5 mg, between 10 pg and 0.4 mg, between 12 pg and 0.3 mg, between 15 pg and 0.2 mg, between 20 pg and 0.1 mg, or between 25 pg and 50 pg. In various embodiments, the kit comprises synthetic peptides at different concentrations. For example, the kit can include a single synthetic peptide at one, two, three, four, five, six, seven, eight, nine, or ten different concentrations. In various embodiments, the kit includes a synthetic peptide at a single concentration that can be further diluted using dilution reagents (dilution reagents which may be included in the kit).

[0121] In various embodiments, a kit comprises a background human proteome comprising one or more human peptides from cells at a quantity of 5 pg, 6 pg, 7 pg, 8 pg, 9 pg, 10 pg, 11 pg, 12 pg, 13 pg, 14 pg, 15 pg, 16 pg, 17 pg, 18 pg, 19 pg, 20 pg, 21 pg, 22 pg, 23 pg, 24 pg, 25 pg, 26 pg, 27 pg, 28 pg, 29 pg, 30 pg, 35 pg, 40 pg, 45 pg, 50 pg, 55 pg, 60 pg, 65 pg, 70 pg, 75 pg, 80 pg, 85 pg, 90 pg, 95 pg, 100 pg, 150 pg, 200 pg, 300 pg, 400 pg, 500 pg, 600 pg, 700 pg, 800 pg, 900 pg, or 1000 pg. In various embodiments, a kit comprises one or more human peptides at a quantity between 1 pg and 1 mg. In various embodiments, a kit comprises one or more human peptides at a quantity between 5 pg and 0.9 mg, between 6 pg and 0.8 mg, between 7 pg and 0.7 mg, between 8 pg and 0.6 mg, between 9 pg and 0.5 mg, between 10 pg and 0.4 mg, between 12 pg and 0.3 mg, between 15 pg and 0.2 mg, between 20 pg and 0.1 mg, or between 25 pg and 50 pg.

[0122] In various embodiments, a kit comprises separate aliquots of synthetic peptides and human peptides from cells. In various embodiments, a kit comprises synthetic peptides at specific concentrations, such that mixing the aliquots of synthetic peptides with the human peptides would achieve a target concentration of the synthetic peptides. In various embodiments, the target concentration of a synthetic peptide is between 0.0001% and 0.1%. In various embodiments, the target concentration of a synthetic peptide when mixed with the human peptides is between 0.0002% and 0.05%, between 0.0003% and 0.04%, between 0.0004% and 0.03%, between 0.00045% and 0.02%, or between 0.0005% and 0.01%. In various embodiments, the target concentration of a synthetic peptide when mixed with the human peptides is 0.0001%, 0.0002%, 0.0003%, 0.0004%, 0.0005%, 0.0006%, 0.0007%,32IPTS / 200244493 10.0008%, 0.0009%, 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.011%, 0.012%, 0.013%, 0.014%. 0.015%, 0.016%, 0.017%, 0.018%, 0.019%, or 0.020%.

[0123] In various embodiments, a kit comprises between 1 and 100 synthetic peptides between 5 and 50 amino acids in length and a background proteome comprising one or more human peptides from cells. In various embodiments, a kit comprises between 1 and 50 synthetic peptides between 5 and 50 amino acids in length and a background proteome comprising one or more human peptides from cells. In various embodiments, a kit comprises between 2 and 30 synthetic peptides between 5 and 50 amino acids in length and a background proteome comprising one or more human peptides from cells. In various embodiments, a kit comprises between 2 and 30 synthetic peptides between 5 and 45 amino acids in length and a background proteome comprising one or more human peptides from cells.

[0124] In various embodiments, kits disclosed herein include protein digestion reagents. Example protein digestion reagents include proteases such as trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin. In various embodiments, kits disclosed herein include one or more protein labeling agents. As one example, a protein labeling agent can be an agent that labels cysteine groups, an example of which is desthiobiotin iodoacetamide (DBIA). In various embodiments, a protein labeling agent are tags that enable multiplex proteomics. Such protein labeling agents can include tandem mass tags (TMTs) such as TMTprol6, TMTpro32, or superheavy TMTpro.

[0125] In various embodiments, kits disclosed herein can include instructions for performing any of the methods disclosed herein, or can include instructions to further access additional instructions (e.g., via a web link) for performing any of the methods disclosed herein.

[0126] For example, kits disclosed herein can include instructions for determining interference in an analytic procedure, instructions for determining interference in mass spectrometry tag-based proteomics, and / or instructions for detecting a processing anomaly, as is disclosed in further detail herein.

[0127] In various embodiments, kits disclosed herein can be stored and / or transported while maintaining the useability of the elements within the kit (e.g., any of the one or more synthetic peptides, the cells, the human peptides, the protein digestion reagents, and / or the one or more protein labeling agents). In various embodiments, kits can include temperature controlling agents to maintain the useability of the elements within the kit. For example, kits33IPTS / 200244493 1can keep the elements in a cold environment (e.g., refrigerated or frozen environment). The cold environment can include reduced temperatures, such as cold temperatures (e.g., between about 20° C. and about 0° C.); and / or freezing temperatures, including for example about 0° C„ -1° C„ -2° C„ -3° C„ -4° C„ -5° C„ -6° C„ -7° C„ -8° C., -9° C„ -10° C„ -12° C„ -14° C„ -15° C., -16° C., -20° C„ -22° C„ -25° C., -28° C„ -30° C„ -35° C., -40° C„ -45° C., -50° C., -60° C., -70° C„ -80° C., -100° C., -120° C., -140° C., -180° C., -190° C., or -200° C. Kits disclosed herein can be stored and / or delivered in a refrigerator, on ice or a frozen gel pack, in a freezer, in a cryogenic freezer, on dry ice, in liquid nitrogen, and / or in a vapor phase equilibrated with liquid nitrogen.

[0128] Kits can include sample containers, which can include any container suitable for storage and or transport of the elements of the kit; such containers including, but not limited to: a cup, a cup with a lid, a tube, a sterile tube, a vacuum tube, a syringe, a bottle, a microscope slide, or any other suitable container. The container can optionally be sterile.Examples

[0129] Details of the disclosed embodiments will be more fully understood from the examples, which are presented herein for illustrative purposes only, and should not be construed as limiting the invention in any way.EXAMPLE 1 : Synthetic Peptides as Standards for Determining Interference in Mass Spectrometry Tag-based Proteomics

[0130] FIG. 3 shows the peptide selection scheme for identifying example synthetic peptides. A total of 27 peptides were identified and synthesized. The 27 peptides are shown in Table 1 as SEQ ID NOs: 1-27. Peptides shown in Table 1 were synthesized by Genscript using standard solid-phase peptide syntheis schemes. Synthetic peptides were purified by reversed phase HPLC to a purity greater than 90%.

[0131] The peptides of SEQ ID NOs: 1-27 were synthesized and assembled into two, linear mini -protein sequences:SEO ID NO: 28:SGRCAALITGHNRDWCLSGKECDIVFSGLDADYAGAIEKIVTLCGDAPEGKRFACQID IPNEENNNASKFTECIDDIKPKFWVNPDCGLKGAAFICAIHSPTLRGFLFPLVETGCNV RLAECDNIGEYIKKEILGTAQSVGCRSEQ ID NO: 29:GSKNCVVLGCERRNVAAGCNPMDLRRQGFGNLPICIAKTVIDPEDTIFIGNVAHECTE DDLKTIPECEAFLKWAAAAVCEKDLPEGCEMKFTGCNVIGLNNNDYQIAKYGICGPN GCGKIYCSGLQCGR34IPTS / 200244493 1

[0132] These sequences were cloned into a bacterial expression vector containing a His-MBP tag, which were expressed in E.coli and purified by standard methods. Purified proteins were then linearized with 5mM Dithiothreitol (DTT) + 2% SDS (sodium dodecyl sulfate) before cysteines were alkylated with desthiobiotin iodoacetamide (DBIA) (lOmM). Purified DBIA labeled proteins were obtained by TCA:Acetone precipitation before resuspension in 1% SDS in phosphate buffered saline to a final concentration of Img / mL.

[0133] Peptide labeling: Lyophilized peptides were resuspended in 50mM EPPS pH8.0 with 20% acetonitrile before disulfides were reduced with TCEP and free cysteines were labeled with DBIA (lOmM, Ihr). Free amines (lysines and peptide n-terminus) were labeled with TMTprol8, TMTpro32, or superheavy TMTpro (Ihr). Labeled peptides were then desalted by solid phase extraction.

[0134] Preparation of Background Peptides of Background Proteome: Expi293 cells were lysed with 50mM EPPS buffer containing 8M Urea and 5mM DTT before cysteines were alkylated with DBIA. Contaminants were removed by SP3, and proteins were digested using Trypsin / LysC. DBIA labeled peptides were enriched with streptavidin resin before labeling with TMTprol8 or TMTpro32.

[0135] Standard preparation: SEQ ID Nos: 1-27 were split into seven pools of three or four peptides each, then titrated into new tubes according to the seven sample KO or dose response schemes (FIG. 4 through FIG. 10). These peptides were then labeled with TMTprol8 or TMTpro32 according to manufacturer specifications. The TMT-labeled peptide pools were desalted by solid phase extraction then mixed with TMT-labeled background proteome. The concentration of each peptide in the final mixure(s) ranged from 3 nanomol per liter (3 nM) to 1 picomol per liter (1 pM).

[0136] MS acquisition: Two mass spectrometry methods were performed to compare precision and sensitivity.

[0137] Results:

[0138] FIGs. 4A and 4B show signal (FIG. 4A) and signal to noise (S / N) (FIG. 4B) values across a plurality of channels for a first knockout sample (KOI). Here in FIG. 4A, the KOI sample includes four peptides (SEQ ID Nos: 1-4), but signal is not expected to be observed in a subset of channels (e.g., knockout channels 1-2, 5-6, 9-10, and 13-14). As shown in FIG. 4B, the corresponding knockout channels 2-3, 6-7, 10-11, and 14-15 include35IPTS / 200244493 1small signal, which is attributable to interference - the IFI for these samples in this acquisition method was calculated as 0.973.

[0139] FIGs. 5A and 5B show signal (FIG. 5A) and signal to noise (S / N) (FIG. 5B) values across a plurality of channels for a second knockout sample (KO2). Here in FIG. 5A, the KO2 sample includes 4 peptides (SEQ ID Nos: 5-8), but signal is not expected to be observed in a subset of channels (e.g., knockout channels 3-4, 7-8, 11-12, and 15-16). As shown in FIG. 5B, the corresponding knockout channels 1, 4-5, 8-9, 12-13, and 16-17 include small signal, which is attributable to interference - the IFI for these samples in this acquisition method was calculated as 0.965.

[0140] FIGs. 6A and 6B show signal (FIG. 6A) and signal to noise (S / N) (FIG. 6B) values across a plurality of channels for a third knockout sample (KO3). Here in FIG. 6A, the KO3 sample includes 3 peptides (SEQ ID Nos: 9-11), but signal is not expected to be observed in a subset of channels (e.g., knockout channels 5-8 and 13-16). As shown in FIG. 6B, the corresponding knockout channels 1, 6-9, and 14-18 include small signal, which is attributable to interference - the IFI for these samples in this acquisition method was calculated as 0.948.

[0141] FIGs. 7A and 7B show signal (FIG. 7A) and signal to noise (S / N) (FIG. 7B) values across a plurality of channels for a fourth knockout sample (KO4). Here in FIG. 7A, the KO4 sample includes 4 peptides (SEQ ID Nos: 12-15), but signal is not expected to be observed in a subset of channels (e.g., knockout channels 1-4 and 9-12). As shown in FIG. 7B, the corresponding knockout channels 1-5 and 10-13 include small signal, which is attributable to interference - the IFI for these samples in this acquisition method was calculated as 0.986.

[0142] FIGs. 8A and 8B show signal (FIG. 8A) and signal to noise (S / N) (FIG. 8B) values across a plurality of channels for a fifth sample across concentrations of a dose response (DR1). Here in FIG. 8A, the fifth sample includes 4 peptides (SEQ ID Nos: 16-19), and signal is expected to increase across the channels (e.g., increasing from channel 1-8 and then from 9-16). As shown in FIG. 8B, the corresponding channels 1-9 and 10-17 exhibit a similar dose-related increase. The slight deviations (e.g., non-linearity) is attributable to interference.

[0143] FIGs. 9A and 9B show signal (FIG. 9A) and signal to noise (S / N) (FIG. 9B) values across a plurality of channels for a sixth sample across concentrations of a dose response (DR2). Here in FIG. 9A, the sixth sample includes 4 peptides (SEQ ID Nos: 20-23), and signal is expected to increase across the channels (e.g., increasing from channel 1-16). As36IPTS / 200244493 1shown in FIG. 9B, the corresponding channels 1-17 exhibit a similar dose-related increase. The slight deviations (e.g., non-linearity) is attributable to interference.

[0144] FIGs. 10A and 10B show signal (FIG. 10A) and signal to noise (S / N) (FIG. 10B) values across a plurality of channels for a seventh sample across concentrations of a dose response (DR3). Here in FIG. 10A, the seventh sample includes 4 peptides (SEQ ID Nos: 24-27), and signal is expected to increase across certain channels (e.g., increasing from channels 1, 3, 5, and 7 and then restarting and increasing from channels 9, 11, 13, and 15). The signal is expected to decrease across certain channels (e.g., decreasing across channels 2, 4, 6, and 8, and then restarting and decreasing from channels 10, 12, 14, and 16). As shown in FIG. 10B, the corresponding channels exhibit similar dose-related changes . The slight deviations (e.g., non-linearity) is attributable to interference.

[0145] FIG. 11 shows background signal across a plurality of channels. Here, the background signal arises from human proteins from the background proteome. The background signal is normalized across the channels.

[0146] Interference as a function of acquisition method: Two acquisition strategies were employed. FIGs. 12A and 12B show no and high interferences across a plurality of channels corresponding to four peptides (SEQ ID Nos: 5-8). Specifically, for the no interference, lower signal / noise (SN) was detected in specific channels (e.g., TMT channels 1, 4, 5, 8, 9, 12, 13, 16, 17 and 18) in comparison to the high interference setting. The low interference and high interference scans are representative of results arising from two different instrument acquisition methods.

[0147] FIGs. 13A and 13B show interference as a function of protein concentration (protein of SEQ ID NO: 26). Specifically, at the higher protein concentration of 1 femtomol (fmol), lower signal / noise (SN) was detected in specific channels (e.g., TMT channels 1, 6, 7, 8, 9, 14, 15, 16, 17 and 18) in comparison to a lower protein concentration at 100 attomole (amol). The low and high interference scans come from the same instrument acquisition method.EXAMPLE 2: Synthetic Peptides for Tracking Processing Anomalies

[0148] This example describes the process of tracking sample preparation artifacts through synthetic peptides. Specifically, FIGs. 14A and 14B show exemplary methods for tracking processing anomalies through engineered proteins. Through the steps shown in FIG. 14B, synthetic proteins are spiked into the sample. As one example, an assembled synthetic37IPTS / 200244493 1protein of SEQ ID NO: 28 or SEQ ID NO: 29 is spiked into a sample following e.g., Step 3: Protein Extraction. Then, Step 4: Trypsin Digest is performed. Subsequently, following mass spectrometry and analysis, any anomalies arising from Step 4: Trypsin Digest are detected.

[0149] FIGs. 15 A and 15B show different signals across a plurality of channels, which shows differences between fully tryptic peptides and missed cleavage peptides. The fully tryptic peptide corresponds to SEQ ID NO: 11 and the missed cleavage peptide corresponds to SEQ ID NO: 12. Here, the differences between fully tryptic peptides and missed cleavage peptides reflect processing anomalies arising from at least Step 4: Trypsin digest, as certain peptides are cleaved and others are missed. Thus, the protein corresponding to the missed cleavage peptide (FIG. 15B) is identified as a protein that experiences processing anomalies, and is thereby useful for informing that a sample is not being fully processed.38IPTS / 200244493 1INCORPORATION BY REFERENCE

[0150] All publications and patents mentioned herein, including those items listed below, are hereby incorporated by reference in their entirety for all purposes as if each individual publication or patent was specifically and individually incorporated by reference. In case of conflict, the present application, including any definitions herein, will control.EQUIVALENTS AND SCOPE

[0151] In the claims articles such as “a,” “an,” and “the” may mean one or more than one unless indicated to the contrary or otherwise evident from the context. Claims or descriptions that include “or” between one or more members of a group are considered satisfied if one, more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process unless indicated to the contrary or otherwise evident from the context. The disclosure includes embodiments in which exactly one member of the group is present in, employed in, or otherwise relevant to a given product or process. The disclosure includes embodiments in which more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process.

[0152] Furthermore, the disclosure encompasses all variations, combinations, and permutations in which one or more limitations, elements, clauses, and descriptive terms from one or more of the listed claims is introduced into another claim. For example, any claim that is dependent on another claim can be modified to include one or more limitations found in any other claim that is dependent on the same base claim. Where elements are presented as lists, e.g., in Markush group format, each subgroup of the elements is also disclosed, and any element(s) can be removed from the group. It should it be understood that, in general, where the disclosure, or aspects of the disclosure, is / are referred to as comprising particular elements and / or features, certain embodiments of the disclosure or aspects of the disclosure consist, or consist essentially of, such elements and / or features. For purposes of simplicity, those embodiments have not been specifically set forth in haec verba herein. It is also noted that the terms “comprising” and “containing” are intended to be open and permits the inclusion of additional elements or steps. Where ranges are given, endpoints are included. Furthermore, unless otherwise indicated or otherwise evident from the context and understanding of one of ordinary skill in the art, values that are expressed as ranges can assume any specific value or sub-range within the stated ranges in different embodiments of39IPTS / 200244493 1the disclosure, to the tenth of the unit of the lower limit of the range, unless the context clearly dictates otherwise.

[0153] This application refers to various issued patents, published patent applications, journal articles, and other publications, all of which are incorporated herein by reference. If there is a conflict between any of the incorporated references and the instant specification, the specification shall control. In addition, any particular embodiment of the present disclosure that falls within the prior art may be explicitly excluded from any one or more of the claims. Because such embodiments are deemed to be known to one of ordinary skill in the art, they may be excluded even if the exclusion is not set forth explicitly herein. Any particular embodiment of the disclosure can be excluded from any claim, for any reason, whether or not related to the existence of prior art.

[0154] Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific embodiments described herein. The scope of the present embodiments described herein is not intended to be limited to the above Description, but rather is as set forth in the appended claims. Those of ordinary skill in the art will appreciate that various changes and modifications to this description may be made without departing from the spirit or scope of the present disclosure, as defined in the following claims.40IPTS / 200244493 1SEQUENCESExemplary sequences of synthetic peptidesAssembled synthetic protein 1 (SEQ ID NO: 28)SGRCAALITGHNRDWCLSGKECDIVFSGLDADYAGAIEKIVTLCGDAPEGKRFACQIDIPNEENNNASKFTECIDDIKPKFWVNPDCGLKGAAFICAIHSPTLRGFLFPLVETGCNVRLAECDNIGEYIKKEILGTAQSVGCRAssembled synthetic protein 2 (SEQ ID NO: 29)GSKNCVVLGCERRNVAAGCNPMDLRRQGFGNLPICIAKTVIDPEDTIFIGNVAHECTEDDLKTIPECEAFLKWAAAAVCEKDLPEGCEMKFTGCNVIGLNNNDYQIAKYGICGPNGCGKIYCSGLQCGR41IPTS / 200244493 1

Claims

CLAIMS1. A method for determining interference and / or sample processing artifacts in mass spectrometry -based proteomics, the method comprising: providing a mixture of one or more synthetic peptides and one or more background peptides of a background proteome, wherein the one or more synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; optionally performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents; performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal.

2. The method of claim 1, wherein the one or more synthetic peptides are between 5 and 45 amino acids in length.

3. The method of claim 1 or 2, wherein the one or more synthetic peptides comprise between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides.

4. A method for determining interference and / or sample processing artifacts in mass spectrometry -based proteomics, the method comprising: providing a mixture comprising between 2 and 30 synthetic peptides between 5 and 45 amino acids in length and one or more background peptides of a background proteome, wherein the one or more synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; optionally performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents; performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal.42IPTS / 200244493 15. The method of any one of claims 1-4, wherein the mixture comprises about 20 synthetic peptides.

6. The method of any one of claims 1-5, wherein the one or more synthetic peptides are selected from a protein comprising a sequence sharing at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or 100% identity to any one of SEQ ID NOs: 1-27.

7. The method of any one of claims 1-5, wherein the mixture comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 1, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, or 27 different synthetic peptides each comprising a protein sequence selected from the group consisting of SEQ ID NOs: 1-27.

8. The method of any one of claims 1-7, further comprising digesting one or more background proteins of a background proteome using protein digestion reagents to generate the one or more background peptides.

9. The method of claim 8, wherein the protein digestion reagents comprise trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin.

10. The method of any one of claims 1-9, wherein the one or more peptide labeling agents comprise tandem mass tags (TMTs).

11. The method of any one of claims 1-10, further comprising: prior to generating the mixture, obtaining the one or more background peptides of the background proteome by: lysing a plurality of cells; and extracting the one or more background peptides from the lysed plurality of cells.

12. The method of claim 11, further comprising performing thiol labeling of the one or more background peptides in the lysed plurality of cells, optionally wherein the thiol labeling comprises contacting the one or more background peptides with desthiobiotin iodoacetamide (DBIA).

13. The method of claim 12, further comprising spiking one or more assembled synthetic proteins into one or more of:(iv) a sample of the lysed plurality of cells;(v) a sample of the thiol-labeled background peptides; and(vi) a sample of the extracted one or more background peptides.

14. The method of claim 12, further comprising spiking one or more assembled synthetic proteins into two or more of:(iv) a sample of the lysed plurality of cells;43IPTS / 200244493 1(v) a sample of the thiol-labeled background peptides; and(vi) a sample of the extracted one or more background peptides.

15. The method of claim 12, further comprising spiking one or more assembled synthetic proteins into each of:(iv) a sample of the lysed plurality of cells;(v) a sample of the thiol-labeled background peptides; and(vi) a sample of the extracted one or more background peptides.

16. The method of any one of claims 13-15, wherein an assembled synthetic protein is assembled from two or more of the synthetic peptides.

17. The method of any one of claims 13-15, wherein an assembled synthetic protein is assembled from two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen synthetic peptides.

18. The method of claim 17, wherein an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14.

19. The method of claim 17, wherein an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27.

20. The method of any one of claims 13-19, wherein determining interference and / or sample processing artifacts across the plurality of channels using the detected signal comprises determining ion interference.

21. The method of any one of claims 1-20, wherein the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%.

22. The method of any one of claims 13-21, wherein the one or more assembled synthetic proteins comprise one or more assembled synthetic, non-naturally occurring peptides.

23. The method of any one of claims 1-22, wherein the one or more synthetic peptides comprise one or more synthetic, non-naturally occurring peptides.

24. The method of any one of claims 1-23, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons,44IPTS / 200244493 1or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice).

25. The method of any one of claims 1-23, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

26. A method for determining interference and / or sample processing artifacts in mass spectrometry-based proteomics, the method comprising: generating a mixture of one or more synthetic peptides and one or more background peptides of a background proteome, wherein concentrations of each synthetic peptide is selected to account for differences in ionization and flight; performing mass spectrometry to detect signal from at least the one or more synthetic peptides, or fragments thereof, across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected signal.

27. The method of claim 26, wherein the one or more synthetic peptides are between 5 and 45 amino acids in length.

28. The method of claim 26 or 27, wherein the one or more synthetic peptides comprise between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides.

29. The method of any one of claims 26-28, wherein the mixture comprises about 20 synthetic peptides.

30. The method of any one of claims 26-29, wherein the one or more synthetic peptides are selected from a protein comprising a sequence sharing at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or 100% identity to any one of SEQ ID NOs: 1-27.

31. The method of any one of claims 26-30, wherein the mixture comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 different synthetic peptides each comprising a protein sequence selected from the group consisting of SEQ ID NOs: 1-27.45IPTS / 200244493 132. The method of any one of claims 26-31, wherein determining interference and / or sample processing artifacts across the plurality of channels using the detected signal comprises determining ion interference.

33. The method of any one of claims 26-32, wherein the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01 %.

34. A method for detecting a processing anomaly, the method comprising: providing a mixture of one or more assembled synthetic peptides and one or more background peptides of a background proteome, the one or more assembled synthetic peptides each assembled from two or more synthetic peptides, and wherein the one or more synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; performing one or more of: digesting the mixture using protein digestion reagents: performing labeling of proteins or fragments thereof in the mixture using one or more peptide labeling agents; and purify ing proteins or fragments thereof in the mixture; performing mass spectrometry to detect signal across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent: and detecting a processing anomaly arising from the one or more of the digestion, labeling, or purifying steps.

35. The method of claim 34, wherein providing the mixture comprises spiking a known concentration of the one or more assembled synthetic proteins into a sample comprising the one or more background peptides of a background proteome.

36. The method of claim 34 or 35, an assembled synthetic protein is assembled from two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen synthetic peptides.

37. The method of claim 36, wherein an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14.46IPTS / 200244493 138. The method of claim 36, wherein an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27.

39. The method of any one of claims 34-38, wherein an assembled synthetic protein comprises a sequence sharing at least 80% identity, at least 85% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.1 % identity, at least 99.2% identity, at least 99.3% identity, at least 99.4% identity, at least 99.5% identity, at least 99.6% identity, at least 99.7% identity, at least 99.8% identity, at least 99.9% identity, or 100% identity with SEQ ID NO: 28 or SEQ ID NO: 29.

40. The method of any one of claims 34-39, wherein the two or more synthetic peptides are between 5 and 45 amino acids in length.

41. The method of any one of claims 34-40, wherein the one or more synthetic peptides comprise between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides.

42. The method of any one of claims 34-39, wherein the protein digestion reagents comprise trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin.

43. The method of any one of claims 34-39, wherein the one or more peptide labeling agents comprise tandem mass tags (TMTs).

44. The method of any one of claims 26-43, wherein the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%.

45. The method of any one of claims 34-44, wherein the one or more assembled synthetic proteins comprise one or more assembled synthetic, non-naturally occurring peptides.

46. The method of any one of claims 34-45, wherein the two or more synthetic peptides comprise two or more synthetic, non-naturally occurring peptides.

47. The method of any one of claims 34-46, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background47IPTS / 200244493 1proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice).

48. The method of any one of claims 34-46, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

49. A kit comprising: one or more synthetic peptides; a background proteome comprising one or more background peptides from cells, wherein the one or more synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; protein digestion reagents; one or more peptide labeling agents; and instructions for determining interference and / or sample processing artifacts in an analytic procedure using the one or more synthetic peptides, the one or more background peptides of the background proteome, the protein digestion reagents, and the one or more peptide labeling agents.

50. The kit of claim 49, wherein the instructions further comprise instructions for: generating, or providing a generated mixture of, the one or more synthetic peptides and the one or more background peptides of the background proteome; optionally performing labeling of proteins or fragments thereof in the digested mixture using the one or more peptide labeling agents; performing a mass spectrometry analytic procedure to detect signal to noise (SN) values across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected SN values.

51. The kit of claim 49 or 50, wherein the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001% to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%.

52. The kit of any one of claims 49-51, wherein the one or more synthetic peptides are between 5 and 45 amino acids in length.48IPTS / 200244493 153. The kit of any one of claims 49-52, wherein the one or more synthetic peptides comprise between 1 and 1000 peptides, optionally wherein the one or more synthetic peptides comprise between 2 and 100 peptides, between 4 and 90 peptides, between 6 and 80 peptides, between 8 and 70 peptides, between 10 and 60 peptides, between 12 and 50 peptides, between 14 and 40 peptides, between 16 and 30 peptides, or between 18 and 22 peptides.

54. The kit of any one of claims 49-53, wherein one or more synthetic peptides comprise two or more synthetic, non-naturally occurring peptides.

55. The kit of any one of claims 49-54, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice).

56. The kit of any one of claims 49-54, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

57. A kit for determining interference and / or sample processing artifacts in mass spectrometry -based proteomics, the method comprising: between 2 and 30 synthetic peptides between 5 and 45 amino acids in length; a background proteome comprising one or more background peptides from cells, wherein the 2 and 30 synthetic peptides differ from the one or more background peptides of the background proteome by at least one amino acid; protein digestion reagents; one or more peptide labeling agents; and instructions for determining interference and / or sample processing artifacts in an analytic procedure using the one or more synthetic peptides, the one or more background peptides of the background proteome, the protein digestion reagents, and the one or more peptide labeling agents.

58. The kit of claim 57, wherein the instmctions further comprise instructions for: generating, or providing a generated mixture of, the one or more synthetic peptides and the one or more background peptides of the background proteome; optionally performing labeling of proteins or fragments thereof in the digested mixture using the one or more peptide labeling agents;49IPTS / 200244493 1performing a mass spectrometry analytic procedure to detect signal to noise (SN) values across a plurality of channels, each channel corresponding to a distinguishable portion of a peptide labeling agent; and determining interference and / or sample processing artifacts across the plurality of channels using the detected SN values.

59. The kit of claim 57 or 58, wherein the generated mixture comprises the one or more synthetic peptides and the one or more background peptides of the background proteome at a ratio between 0.0001 % to 0.05%, optionally at a ratio between 0.0002% to 0.04%, between 0.0003% to 0.03%, between 0.0004% to 0.02% or between 0.0005% to 0.01%.

60. The kit of any one of claims 49-59, wherein the mixture comprises about 20 synthetic peptides.

61. The kit of any one of claims 49-60, wherein the one or more synthetic peptides are selected from a protein comprising a sequence sharing at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or 100% identity to any one of SEQ ID NOs: 1-27.

62. The kit of any one of claims 49-60, wherein the mixture comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 different synthetic peptides each comprising a protein sequence selected from the group consisting of SEQ ID NOs: 1-27.

63. The kit of any one of claims 49-62, wherein the protein digestion reagents comprise trypsin, Glu-C, LysN, Lys-C, Asp-N, or chymotrypsin.

64. The kit of any one of claims 49-63, wherein the one or more peptide labeling agents comprise tandem mass tags (TMTs).

65. The kit of any one of claims 49-64, wherein the kit further comprises an assembled synthetic protein assembled from two or more of the synthetic peptides.

66. The kit of claim 65, wherein an assembled synthetic protein is assembled from two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen synthetic peptides.

67. The kit of claim 65, wherein an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, or fourteen sequences shown as SEQ ID NOs: 1-14.

68. The kit of claim 65, wherein an assembled synthetic protein is assembled from synthetic peptides comprising two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, or thirteen sequences shown as SEQ ID NOs: 15-27.50IPTS / 200244493 169. The kit of any one of claims 57-68, wherein the 2 to 30 synthetic peptides comprise 2 to 30 synthetic, non-naturally occurring peptides.

70. The kit of any one of claims 57-69, wherein n the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice).

71. The kit of any one of claims 57-69, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

72. A mixture for determining analytic interference, the mixture comprising: between 1 to 20 synthetic peptides each comprising a sequence selected from the group consisting of SEQ ID NOs: 1-27.

73. The mixture of claim 72, wherein the mixture further comprises one or more background peptides of a background proteome.

74. The mixture of claim 73, wherein the one or more background peptides of the background proteome are obtained by: lysing a plurality of cells; and extracting the one or more background peptides from the lysed plurality of cells.

75. The mixture of any one of claims 73-74, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from any of a unicellular organism (e.g., bacteria or yeast), or multicellular organism (e.g., humans, non-human primates (such as orangutans, baboons, or chimpanzees), horses, cows, pigs, sheep, goats, dogs, Cats, rabbits, guinea pigs, gerbils, hamsters, rats and mice).

76. The mixture of any one of claims 73-74, wherein the one or more background peptides of a background proteome comprise one or more background peptides of a background proteome obtained from one or more humans.

77. The mixture of any one of claims 72-76, wherein the 1 to 20 synthetic peptides comprise 1 to 20 synthetic, non-naturally occurring peptides.51IPTS / 200244493 1