A novel method for sample preparation for qualitative and quantitative analysis of the serum peptidome based on amino acid sequencing using tandem mass spectrometry, and a novel data analysis workflow for quantitative analysis of peptides sequenced using tandem mass spectrometry.
The novel serum peptidomic method addresses the limitations of existing analysis by dissociating peptides from high-abundance proteins using a citrate-phosphate buffer and hydrophilic-lipophilic balance column, enabling the high-throughput identification and quantification of a large number of serum peptides.
Patent Information
- Application Number
- JP2025536504
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-23
- Filing Date
- 2023-11-08
- Publication Date
- 2026-01-27
AI Technical Summary
Existing methods for serum peptidome analysis are limited by the complexity of serum, high-abundance proteins masking low-abundance peptides, instability of low-abundance proteins, and the need for costly and time-consuming depletion and purification steps, resulting in a low number of identified peptide biomarkers.
A novel sample preparation method involving incubation with a citrate-phosphate buffer at pH 3.1-3.6 to dissociate peptides from high-abundance proteins, followed by purification using a hydrophilic-lipophilic balance column, eliminating the need for depletion and precipitation steps, and enzymatic digestion, enabling direct sequencing and quantification of natural peptides.
This method allows for the high-throughput identification and quantification of 15,000-17,000 serum peptides, including 1,900-2,100 quantified peptides, without the need for expensive kits or complex procedures, facilitating comprehensive qualitative and quantitative peptidomic analysis.
Smart Images

Figure 2026502865000001 
Figure 2026502865000002 
Figure 2026502865000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to serum sample preparation, qualitative and quantitative methods for serum peptidome profiling of human samples. The approach of the present invention is simple, cost-effective, rapid and comprehensive for serum peptidomic profiling. These serum peptides can be used for biomarker discovery, diagnosis, prognosis, prediction, monitoring and differentiation of various human disease types. [Background technology]
[0002] Peptidomics targets and analyzes peptides (by their length, distribution, protein origin, or de novo sequencing). Peptides are protein fragments with diverse functions. Because peptides regulate multiple physiological and pathological processes, changes in the peptidome can reflect both beneficial and harmful phenomena. Therefore, peptidome analysis of biological fluids, such as serum samples, is a promising approach for identifying peptide biomarkers of various conditions and monitoring patient response to applied treatments. The peptidome [1] is considered to be low-molecular-weight peptides (less than 15 kDa). Endogenous peptides primarily act as messengers (e.g., cytokines, hormones, growth factors) and signal various biological processes. Therefore, these peptides can reflect a person's health or disease state. In addition, endogenous peptides also provide insight into proteolytic activity, degradation, and denaturation, enabling the study of the disease microenvironment.
[0003] Despite the introduction of peptidomics over the past decade, serum peptidome analysis has consistently been limited for a variety of reasons: a) serum is highly complex, containing components of all proteins produced in the body; b) high-abundance proteins, such as serum albumins and immunoglobulins, account for nearly 90% of total serum protein by weight, thus obscuring the detection of other proteins, especially low-abundance proteins or peptides; c) simultaneously limiting the detection of potentially diagnostic low-abundance peptides; d) the disproportionate concentration of low-abundance and high-abundance proteins in the total serum protein concentration (high-abundance proteins account for around 95% and low-abundance proteins account for less than 5%); e) the instability of low-abundance proteins or peptides (hormones or low-abundance proteins or potential diagnostic or prognostic biomarkers); and f) the overwhelming number of depletion steps, fractionation, and precipitation methods to remove high-abundance proteins reduces the number of identified diagnostic peptides and the yield, potentially losing important biological information to background noise. Thus, all previously described methods for serum peptide analysis were characterized by a low number of identified peptides, and no quantitative methods for serum peptidome analysis were introduced.
[0004] The present invention involves unique sample preparation steps for serum peptidomic qualitative and quantitative analysis. The use of serum for early screening and diagnosis is essential for developing effective treatments for many diseases. Data from previous literature has shown that serum peptidomics approaches have been used to diagnose a wide range of diseases, such as a lung cancer study involving 64 patients and identifying only 38 potential peptide biomarkers to distinguish lung cancer patients from healthy controls [2]; a reproductive disorder study involving 24 patients and showing, after statistical analysis, that 23 peaks / peptides on average had significantly different peak intensities in pregnancy-induced hypertension (PIH) patients compared to controls [3]; a gestational diabetes mellitus study involving 200 patients and showing that a total of 297 identified peptides were significantly differentially expressed in the gestational diabetes mellitus (GDM) group compared to controls [4]; a rheumatoid arthritis study involving 35 patients and showing 113 peaks that distinguished RA from primary osteoarthritis (OA) and 101 peaks that distinguished RA from healthy controls (HC) (although no quantification was performed) [5]; and an ulcerative colitis study involving 78 patients and showing that 732 peptides were successfully identified by this method [6]. However, the ultimate success of these previously published methods is severely limited by the small number of identified markers. There is a need to develop simple and effective serum sample preparation methods that would enable the quantification of serum peptidomic data. Known methods either identify a low number of peptides or quantify only a small number of peptides (fewer than 20 peptides). Complex sample preparation, including depletion of high-abundance proteins, precipitation, lysis, protein labeling with fluorescent dyes, multi-step peptide purification, and desalting steps, requires additional chromatographic setup. Our approach, however, bypasses all these time-consuming and costly steps.
[0005] In the past, several patents were published using methods, devices, and kits aimed at addressing biomarker discovery using different approaches, such as protein / peptide levels and immunoassays [7][8][9]
[10]
[11] , plasma enzyme levels
[12] , etc. Several patents
[11] reported one or more peptide biomarkers for diagnosing cardiovascular disease (CVD) by analysis with mass spectrometry (MS), such as matrix-assisted laser desorption / ionization (MALDI)
[11] and time-of-flight (TOF)
[12] . Protein quantification data were missing in these reports [8][9]. A total of 20 patient serum samples from
[11] were analyzed for free and bound peptides of total serum proteins. Furthermore, when considering single biomarkers such as glycogen phosphorylase-BB (GPBB) [7] or protein alpha-1 antitrypsin
[13] , L-glutamine hydroxylamine glutamyltransferase (L-GHGT) and gamma-glutamylhydroxamate synthetase (GGHS) activity
[12] in stroke, the risk is too great to rely on for accurate diagnosis. In contrast, our approach is high-throughput and comprehensive, identifying 15,000–16,000 serum peptides in 15 patients and successfully quantifying 1,900–2,100 serum peptides. Summary of the Invention
[0006] The present inventors have created a serum peptidomic approach (based on amino acid sequences) that differs from the available approaches. It is based purely on the sequencing of natural peptides of amino acid sequences, but there are significant differences from existing methods, such as sample preparation, time, and cost aspects, in proteomics / peptidomic approaches for the diagnosis and prognosis of cardiovascular diseases [9]
[10]
[11] . Our novel peptidomic approach is unique in many steps. 1. The main difference is that it focuses on serum peptides rather than proteins. 2. A unique step involves concentrating peptides in serum samples by incubating them with a citrate-phosphate buffer solution of pH 3.1-3.6 in water of 99.9% purity or higher, where the volume ratio of buffer to serum is 29:1, and mixing the buffer and serum at a temperature ranging from 4 to 8°C to dissociate peptides from high-abundance serum proteins. A hydrophilic-lipophilic balance column is used to remove proteins of 15 kDa or larger from the concentrated peptides, while simultaneously purifying the peptides to obtain peptides of less than 15 kDa. Therefore, it features the unexpected use of a low-pH buffer to concentrate peptides from complexes with high-abundance / carrier proteins. 3. In the past, they utilized depletion steps, removing high-abundance serum proteins by precipitation [9]
[14] [6]. It is now clear that depleting these high-abundance proteins may exclude unique peptide biomarkers, resulting in poor prognosis / diagnosis with existing biomarkers
[15] . Therefore, an approach without the depletion and precipitation steps was developed, which allows for the generation of more peptides and the isolation of potential carrier peptides from high-abundance proteins. 4. The present invention does not require the use of purification kits such as Ettan 2D clean or any quantification kit protein [8][9], all of these steps are time-consuming and expensive. 5. Previous patents [9] used fluorescent dyes (Cy3, Cy5) for protein labeling due to their high sensitivity. Reading and imaging proteome profiles from dye-based gels requires specific equipment (Typhoon instrument, GE Healthcare, required variable mode imager to generate digital images of radioactive, fluorescent, or chemiluminescent samples), which also requires investment of time and money. Therefore, this is a complex sample preparation procedure and requires experienced personnel to perform this analysis. 6. The sixth and most important difference is the enzymatic digestion step of proteins. To date, no proteomic analysis can be performed without enzymatic digestion (trypsin, chymotrypsin, c-lysine, etc.). In contrast, this approach does not involve any additional manipulation steps such as reduction, alkylation, and enzymatic (trypsin) digestion steps. It also allows for the identification of natural / endogenous peptides (which can act as unique biomarkers) without any modifications. 7. The present invention does not require any lysis buffer to isolate and desaturate the proteins or peptides of interest and keep them in a stable environment. 8. Most reports indicate that they are focused only on protein / peptide identification, but the present invention is based on the performance of both (qualitative and quantitative) analyses. 9. Regarding the equipment for peptide separation and identification (we used the state-of-the-art mass spectrometry model, the Exploris 480 MS), we have previously used platforms based on two-dimensional difference gel electrophoresis (2D DIGE), surface-enhanced laser desorption / ionization (SELDI), or modified liquid chromatography-matrix-assisted laser desorption / ionization mass spectrometry (LC-MALDI), and matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI TOF / MS) for various biological samples. These platforms are less sensitive and comprehensive than orbitrap mass spectrometers. In short, the method of the present invention was developed as a novel and straightforward qualitative, quantitative, and comprehensive approach to serum peptidomics without precipitation, depletion, lysis, protein labeling with fluorescent dyes, gel electrophoresis, desalting, etc. 10. The simplicity and small number of steps to perform during sample preparation makes it easy to perform high-throughput sample preparation (100-200 samples / day). 11. Because serum is the most commonly available sample worldwide and is routinely used for pathological testing, our approach is broadly applicable to all serum samples (human and non-human serum samples, e.g., dog, cat, pig, horse, cow, cheetah, leopard, etc.). 12. This multi-biomarker peptide panel can be easily deployed for various purposes, such as early diagnosis, prognosis, monitoring, and prediction of health problems. Furthermore, it can be used to develop various point-of-care diagnostic (POCD) platforms for various diseases. 13. The method is easily adaptable and, if desired, can enrich for short (<3 kilodaltons / kDa) and long (<10 kilodaltons / kDa) serum peptide length distributions from clinical samples. 14. The impact of our innovation also has enormous potential to translate to serum peptide biomarkers, commercialization, and direct biomarker discovery.
[0007] The present invention aimed to develop a comprehensive sample preparation and method for qualitative and quantitative analysis of serum peptidome samples. Furthermore, this novel approach can be routinely used to discover multiple biomarkers for diagnosing, prognosing, monitoring, and predicting various diseases. This approach facilitates both qualitative and quantitative analysis of serum peptidomics, enabling the identification of a total of 15,000–17,000 serum peptides primarily through data-dependent acquisition (DDA) (iProphet iprob value >0.99%, 15 samples). At the same time, 1,900–2,100 serum peptides were successfully quantified through 30 LC-MS / MS (liquid chromatography with tandem mass spectrometry) analyses of 15 clinical samples through data-independent acquisition (DIA). It required very small amounts of serum sample. Therefore, it is feasible to perform peptidomic analysis using a 100 μL volume of each sample. DETAILED DESCRIPTION OF THE INVENTION
[0008] ESSENCE OF THE INVENTION - DESCRIPTION OF THE INVENTION The novel analytical method for serum peptidomics included the following steps. According to the invention, the method is characterized in that it comprises the following steps: I. Preparation of test serum samples carrying out the following steps: 1. Incubate the test serum with a pH 3.1-3.6 citrate-phosphate buffer solution prepared with ultrapure water (≥99.9%) of liquid chromatography-mass spectrometry (LC-MS) grade, at a volume of 2900 μl for every 100 μl of serum sample, using a 29:1 volume / volume (v / v) ratio of buffer to serum sample. The incubation step takes 3-5 minutes with mixing at 4-8°C to allow dissociation of peptides from highly concentrated serum proteins (e.g., albumin, globulin, transferrin, haptoglobin, alpha-1-antitrypsin, etc.).
[0009] 2. High-abundance proteins ranging from 16 to 170 kDa, such as albumin, globulin, transferrin, haptoglobin, and alpha-1-antitrypsin, mask low-abundance serum peptides less than 15 kDa in test serum samples. These high-abundance proteins interfere with the low-abundance serum peptides less than 15 kDa, preventing their identification. Therefore, the present invention enables the separation and purification of serum biomarker peptides from high-abundance serum proteins in a single step.
[0010] The isolated peptides are purified using a hydrophilic-lipophilic balance column, which removes proteins larger than 15 kDa from the enriched peptides, while simultaneously purifying the peptides using a hydrophilic-lipophilic balance column to obtain peptides smaller than 15 kDa. Therefore, the present invention does not include any additional steps for depletion and precipitation of high-abundance / high-concentration serum proteins, ultimately improving the final yield of serum peptides and preventing the loss of many biomarker peptides during the depletion and precipitation steps. Before loading the incubated serum sample, the hydrophilic-lipophilic balance column is equilibrated with liquid chromatography-mass spectrometry (LC-MS)-grade water and activated with LC-MS-grade methanol / 0.2 volume / volume (v / v) formic acid with a purity of 99.9% or higher.
[0011] 3. The next step is peptide elution, which involves washing the column three to five times with LC / MS-grade water of 99.9% purity or higher to remove unbound and low-binding peptides. The bound serum peptides are then eluted from the hydrophilic-lipophilic balance column using water / 60-80% volume / volume (v / v) methanol / 0.1-0.2% volume / volume (v / v) formic acid. Peptides with molecular weights of 3-10 kilodaltons (kDa) or less are considered to be relevant serum peptides. Therefore, the eluted peptides in the above mixture are passed through 3-10 kilodalton (kDa) filters, which retain peptides with molecular weights greater than 3-10 kilodaltons (kDa). Finally, all eluted serum peptides are lyophilized to dryness. The present invention does not require any cell disruption buffers for peptide isolation, further reduction or alkylation steps, or enzymatic digestion with trypsin or chymotrypsin to maintain the peptides in a stable environment.
[0012] 4. Based on all the above three steps, the method of the present invention is cost-effective, easy to perform, and relatively rapid for serum peptidomic sample preparation, thus facilitating the task of high-throughput peptidomic sample preparation.
[0013] 5) Serum is a commonly available sample and is routinely used for pathological testing worldwide. Therefore, the invented methodology for serum peptidomics is widely applicable to all serum samples from human and non-human species, such as dogs, cats, pigs, horses, cows, cheetahs, and leopards, for serum peptidomic analysis.
[0014] II. Analysis of selected serum peptides 6. The isolated peptides are sequenced based on their amino acid sequence using tandem mass spectrometry tools.
[0015] a) Comprehensive qualitative and quantitative This invention utilizes a multi-search engine strategy to create a spectral library for qualitative and quantitative serum peptidomics, facilitating discrimination between test conditions. As a result, comprehensive qualitative serum peptides are identified for the first time. Furthermore, it identifies a total of 15,000–17,000 serum peptides by data-dependent acquisition (DDA) (iprophet iprobability >0.99%, 15 samples). Simultaneously, 1,900–2,100 serum peptides were successfully quantified by 30 LC-MS / MS (liquid chromatography with tandem mass spectrometry) analyses of 15 clinical samples by data-independent acquisition (DIA). Therefore, high throughput in terms of qualitative and quantitative serum peptidomics is key to this novel serum peptidomics invention.
[0016] b) Peptide length distribution The present invention facilitates a key feature of serum peptidomics by screening a wide range of peptide lengths. These peptides are not generated by trypsin digestion, in which the trypsin enzyme cleaves proteins at every K (lysine) or R (arginine) site. In contrast, the novel approach allows for screening of natural peptides without including any enzymatically digested peptides. Furthermore, the peptide length distribution is broad, ranging from 7 to 30 amino acid residues. Therefore, the range of natural serum peptides from stroke samples can be easily adjusted to 1 to 10 kilodaltons.
[0017] c) Data analysis workflow This invention deals with high-throughput (1900-2100 peptides quantified) quantitative analysis of serum peptidomics. The data analysis workflow for quantitative analysis of peptides sequenced by tandem mass spectrometry included the following steps within the Skyline MSStats pipeline: a) MS / MS DDA (data-dependent acquisition) data were searched with two or more search engine strategies to generate a comprehensive spectral library. b) Quantitative peptidomic data (independent MS / MS data, DIA) were extracted using a comprehensive spectral library. c) Extraction of product ion chromatograms and integration of product ion peak areas (area under the curve, AUC). d) The false discovery rate (FDR) of quantitative extraction / analysis was controlled by determining the FDR values of peak groups in the mProphet module. e) Inference of product ion peak areas for peptides and their sum peak area statistics in the MSstats module was performed from a peptide-centric perspective. f) Further downstream analysis, such as graphical data visualization including volcano plots, dendrograms, heat maps, and functional stroke serum peptide enrichment maps, was performed using quantitative matrices from MSstats. Thus, downstream analysis is not a limitation of the method.
[0018] d) Unsupervised correlation matrix clustering based on quantitative serum peptides of the groups The method of the present invention for serum peptidomics allows for the differentiation of three test conditions. In particular, peptide heat maps and unsupervised hierarchical clustering among the three groups provide an overview of quantitative peptidomic data analysis among the test conditions of a serum peptidome DIA-MS run (Data Independent Acquisition-Mass Spectrometry). The heat maps show similar peptide intensity patterns among the compared groups, resulting in complete clustering within the groups. Differential clustering of serum peptidome groups was observed. Surprisingly, similar log2 peptide intensity patterns were observed among the compared groups, resulting in complete clustering of the control and each test sample.
[0019] e) Multi-bioinformatics approach After the comprehensive quantitative matrix was obtained, the present invention enabled downstream analysis of various multi-bioinformatics data interpretations. These multi-bioinformatics analyses revealed functional understanding. These included GO enrichment analyses (molecular function, biological process, and component analysis) using the Search Tool for the Retrieval of Interacting Genes / Proteins (STRING), Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, REACTOME, The Database for Annotation, Visualization, and Integrated Discovery (DAVID), and keyword enrichment (Uniport). First, we created an interactome map using STRING to identify characteristic node-enriched control and test groups and simultaneously identify other potential serum peptide biomarkers.
[0020] Detailed Nature of the Invention and Research Description - The approach that led to the present invention is shown in FIG.
[0021] The main parts of the scheme of the present invention are: 1. serum peptidomics sample preparation (the single dagger symbol † indicates a novel step using an acidic buffer (pH 3.1-3.6) for peptide enrichment; the double dagger symbol (††) indicates another novel single step of depleting / removing high-abundance proteins and purifying serum peptides, i.e., using a hydrophilic-lipophilic balance column to remove proteins larger than 15 kDa from the enriched peptides and simultaneously purifying the peptides to obtain peptides smaller than 15 kDa); 2. sequencing of serum peptides by tandem mass spectrometry; and 3. qualitative and quantitative data analysis (the triple dagger symbol ††† indicates a novel combined step of steps 3.1-3.4 for comprehensive qualitative and quantitative serum peptidomics).
[0022] The present invention, unlike those available approaches, has created a serum peptidomic approach (based on amino acid sequences), which is based purely on the sequencing of natural peptides of amino acid sequence and includes the following steps:
[0023] A quick explanation of how 1. Samples Collection of serum samples from patients / disease groups and respective control groups. A 50-100 μl serum sample is collected from each patient and immediately stored at -80°C until analysis.
[0024] 2. Quality of Glassware Mass spectrometers are highly sensitive analytical instruments that can detect even minute amounts of contaminants during sample preparation. To avoid compromising peptide signals, it is necessary to use highly pure water and other chemicals.
[0025] 3. Buffer Preparation. The unexpected use of low pH buffers for peptide concentration and release from complexes with high abundance / carrier proteins. Citrate-phosphate buffers pH 3.1-3.6 were prepared with high-grade water.
[0026] 4. Incubation: Peptidomics sample preparation begins by incubating serum samples in a glass beaker with an acidic citrate-phosphate buffer (pH 3.1-3.6) at 4-8°C for 3-5 minutes (gentle manual vortexing required) to dissociate protein-peptide interactions, especially peptides bound to carrier proteins such as albumin. Therefore, peptide enrichment and release from complexes with high-abundance / carrier proteins are characterized by the unexpected use of a low pH buffer.
[0027] 5. Peptide Purification: A single-step purification was performed using a hydrophilic-lipophilic balance column. Unexpectedly, this column was able to deplete high-abundance proteins without peptide loss, while simultaneously purifying peptides from other contaminants in serum. This means that the hydrophilic-lipophilic balance column was used to remove proteins larger than 15 kDa from the enriched peptides and simultaneously purify the peptides to obtain peptides smaller than 15 kDa, thereby producing more peptides and enabling the isolation of potential carrier peptides from high-abundance proteins. Before loading the incubated serum sample, the column was equilibrated with water and activated with high-grade methanol.
[0028] 6. Peptide Elution The column was first washed several times with high-grade water to remove unbound and low-bound peptides. The bound serum peptides were then eluted from the column using water / 80% v / v methanol / 0.2% v / v formic acid. The collected additional eluted peptides were passed through a 3 kilodalton molecular weight filter. Finally, the corresponding peptides were lyophilized to dryness. All dried peptide samples were stored at -80°C until mass spectrometry analysis.
[0029] 7. Peptide Sequencing with Tandem Mass Spectrometry Dried serum peptides were further subjected to sequencing on an UltiMate™ 3000 RSLCnano liquid chromatograph (Thermo Scientific) online coupled to an Orbitrap Exploris 480 mass spectrometer (MS) (Thermo Scientific). Indexed retention time (iRT) peptides were added to each sample to control the retention time of serum peptides for downstream quantitative analysis. Separation of analytical peptides was performed by a nonlinear gradient using mobile phases A and B. Serum peptides eluting from the analytical column were ionized with a nanoelectrospray ion source (NSI) connected to a mass spectrometer.
[0030] 8. Data-Dependent Acquisition (DDA) Mass Spectrometry Measurements. Data-dependent acquisition (DDA) was developed for serum peptidomics. A full scan was followed by fragmentation of the top 20 most intense precursor ions and acquisition of their MS / MS spectra. DDA data were acquired at a resolution of 120,000, utilizing a precursor range of m / z 350–1650, and only precursor charge states between +1 and +6 were included in the experiment.
[0031] 9. Data-Independent Acquisition (DIA) Mass Spectrometry Measurements. Each DIA (data-independent acquisition) cycle was performed with the acquisition of 62 precursor windows / scan events while maintaining the same LC (liquid chromatography) separation parameters. The Orbitrap Exploris 480 mass spectrometer was operated in positive polarity data-independent mode (DIA) with a full scan in profile mode at 60,000 resolution. The normalized collision energy and AGC (automatic gain control) were optimized for DIA measurements to provide maximum injection time.
[0032] 10. Data Analysis. The present invention involves comprehensive, qualitative, and quantitative data analysis of serum peptidomics in cardiovascular disease, such as stroke patient populations. Each sample was measured in one DDA and two DIA runs. These technical replicates were used to ensure the highest LC-MS / MS (liquid chromatography with tandem mass spectrometry) assay quality.
[0033] 11. Multi-Search Engine Strategy. Each search engine has its own built-in algorithm to consider the comprehensiveness of data analysis. Therefore, the present invention utilized a multi-search engine strategy for creating spectral libraries. For qualitative analysis of serum peptidomes, DDA (data-dependent acquisition) and DIA (data-independent acquisition) raw files were converted to their respective formats and then searched for specific compatible search engines, such as Comet and Xtandem / MS fragger. Data searches against a human database were coupled with indexed retention time (iRT) peptide sequences (Biognosys), decoy-inverted target sequences, and contaminants. Various search parameters and settings were used for each search engine to construct the spectral library.
[0034] 12. Statistical Filtration. Further recalculation of peptide probabilities was included in downstream analysis. The resulting multi-engine search files were processed with PeptideProphet and iProphet. Recalculated pep.XML files with peptide iprobabilities higher than 0.99 (1% FDR) were then further processed in Skyline to automatically generate retention time calculators based on the indexed retention time peptide retention time values observed in the recalculated pep.XML files. Further serum peptidomic matrices were imported into MSStats for statistical evaluation, including normalization, imputation, t-tests, fold-change calculations, and adjusted p-values.
[0035] 13. Quantitative Peptidomics Data Extraction Continuing with the Skyline file saved in the previous step, the extraction settings were appropriately configured. Furthermore, the DDA (data-dependent acquisition) data was used to generate a spectral library that would not only be used for qualitative serum peptidomics but would also later be required for extracting quantitative data from DIA. The present invention considers DIA (data-independent acquisition) data to be the only option for adequately quantifying serum peptides due to its fully reproducible multiplexing. Furthermore, as previously suggested, DIA (data-independent acquisition) data can be used for both identification and precise quantification of serum peptides. Furthermore, quantitative values reflecting the peptide abundance in the sample (peptide peak area) were extracted from the DIA (data-independent acquisition) data based on the peptide-product ion pairs (transitions) listed in the spectral library from the previous step.
[0036] 14. The present invention provides peptide quantification between compared patient groups to screen for significantly dysregulated peptides. The present invention also considers the adjusted P value (adj.pval), which is the minimum family-wise significance level at which a particular comparison is declared statistically significant as part of multiple comparisons, such as between a test group and its respective control group. Next, fold changes are calculated by relative comparison of the integrated peak areas of the compared conditions. A fold change (fold change > 2) reflects the ratio of the integrated peak areas in the compared conditions and indicates how many times more or less a peptide is present in the first condition compared to the second condition. Therefore, these total quantified peptides can be used as biomarkers to distinguish them from their respective control groups based on their increased levels determined from the corresponding MS / MS (tandem mass spectrometry) signal intensities. The data analysis workflow for quantitative analysis of peptides sequenced by tandem mass spectrometry included the following steps within the Skyline MSStats pipeline: a) MS / MS DDA (data-dependent acquisition) data were searched with two or more search engine strategies to generate a comprehensive spectral library. b) Quantitative peptidomic data (independent MS / MS data, DIA) were extracted using a comprehensive spectral library. c) Extraction of product ion chromatograms and integration of product ion peak areas (area under the curve, AUC). d) The false discovery rate (FDR) of quantitative extraction / analysis was controlled by determining the FDR values of peak groups in the mProphet module. e) Inference of product ion peak areas for peptides and their sum peak area statistics in the MSstats module was performed from a peptide-centric perspective. f) Further downstream analysis, such as graphical data visualization including volcano plots, dendrograms, heat maps, and functional stroke serum peptide enrichment maps, was performed using quantitative matrices from MSstats. Thus, downstream analysis is not a limitation of the method.
[0037] Furthermore, it can be used to develop various point-of-care diagnostic (POCD) platforms for targeted diseases. Therefore, the impact of our innovation also has great potential for translation, commercialization, and discovery of serum peptide biomarkers. This innovation can open up various fields for deadly diseases based on peptide profiling and can grow in academic, medical, biotechnology, and industrial development.
[0038] example A novel qualitative and quantitative serum peptidomics methodology is applied to human and non-human serum samples. The methodology is not limited to serum samples from stroke or healthy donors. To eliminate the randomness of sample batches, 15 clinical samples, i.e., 15 samples from each group, were analyzed. 1. Serum from healthy donors (n=5), 2. Acute ischemic stroke (AIS) routinely diagnosed by imaging modalities computed tomography (CT) or magnetic resonance imaging (MRI) (n=5); 3. Patient samples (n=5) with intracranial hemorrhagic stroke (ICH) routinely diagnosed by imaging modalities computed tomography (CT) or magnetic resonance imaging (MRI). Five samples were examined.
[0039] The present invention is described in more depth in the examples and illustrated in the figures, where in Figure 1 the method of peptidomics sample preparation is shown, in Figure 2 the Venn diagram of qualitative analysis comparing identified peptides is shown, in Figure 3 the developed DIA (Data Independent Acquisition) method is checked against the performance of classical DDA (Data Dependent Acquisition), in Figure 4 the Venn diagram comparing serum peptide hits among three groups (healthy volunteers, acute ischemic stroke (AIS) and intracranial hemorrhage (ICH) patient samples), in Figure 5 the peptide length distribution of serum peptidomics, and in Figure 6 the DDA (Data Dependent Acquisition)-quality analysis. Figure 7 shows 30 individual correlation heatmaps from quantitative analysis (MS). Figure 7 shows peptide heatmaps and unsupervised hierarchical clustering among three groups (healthy donor, acute ischemic stroke (AIS), and intracranial hemorrhage (ICH) patient samples). Figure 8 shows peptide volcano plots of dysregulated serum peptides among three groups (healthy donor, acute ischemic stroke (AIS), and intracranial hemorrhage (ICH) patient samples). Figures 9 and 10 show STRING analysis for acute ischemic stroke (AIS) and intracranial hemorrhage (ICH) samples determined from quantitative comparisons. A novel, high-throughput serum peptidomics approach was used to demonstrate the concept of serum peptidomics analysis. Samples were collected, serum was isolated within 1 hour, and stored at -80°C until sample preparation. Figure 11 shows a detailed description of the nature and research of the present invention—the approach leading to the present invention. [Brief explanation of the drawings]
[0040] [Figure 1] Three-step flowchart of serum peptidomics sample preparation. The first step utilizes a 3-5 minute acidic buffer incubation to allow dissociation of protein-peptide interactions and enrichment of peptides. The second critical step is purification using an Oasis column. [Figure 2] Venn diagram of qualitative analysis comparing peptides identified by DDA (data-dependent acquisition) and DIA (data-independent acquisition) spectroscopy-based peptidomic research methods in all peptidome samples. [Figure 3] Performance comparison of the developed DIA (data-independent acquisition) method with classical DDA (data-dependent acquisition). (Control / CON, acute ischemic stroke / AIS, intracranial hemorrhagic stroke (ICH)). Bar graphs highlight the contribution of the four exploration approaches to the overall identification. [Figure 4] Venn diagram comparing serum peptide hits among three groups (acute ischemic stroke / AIS, intracranial hemorrhagic stroke / ICH, and control / CON). The relatively small overlap may again correspond to patient heterogeneity, but in this case may primarily correspond to different components of patient sera in different patient conditions. [Figure 5] Peptide length distribution in serum peptidomics. [Figure 6] Correlation heatmap of MS (mass spectrometry) runs. Thirty individual DIA-MS (data-independent acquisition-mass spectrometry) runs of the serum peptidome were ordered according to the group compared (CON / healthy controls [n=5], acute ischemic stroke / AIS stroke [n=5], intracranial hemorrhagic stroke / ICH stroke [n=5]). Each patient serum sample was analyzed in technical duplicates. Squares in the heatmap indicate correlations of log2-transformed peptide intensities. [Figure 7]Peptide heatmap and unsupervised hierarchical clustering of 30 individual stroke patient and healthy donor serum peptidome DIA-MS (Data Independent Acquisition-Mass Spectrometry) runs. Heatmap squares represent log2-transformed peptide intensities. The color scale ranges from white (lowest peptide intensity) to black (highest peptide intensity). As expected, clustering across technical replicates demonstrates the superior performance of the LC-MS / MS (liquid chromatography with tandem mass spectrometry) system. The heatmap shows similar peptide intensity patterns among the compared groups, with complete clustering between the healthy donor serum peptidome and the stroke group serum peptidome. Surprisingly, distinct clustering is observed between the serum peptidomes of intracranial hemorrhagic stroke (ICH) and acute ischemic stroke (AIS). [Figure 8A] Peptide volcano plot shows (as dark dots) A. Dysregulated serum peptides that were significant in acute ischemic stroke / AIS vs. healthy controls. [Figure 8B] Peptide volcano plot shows (as dark dots) B. Dysregulated serum peptides that were significant in acute ischemic stroke / AIS vs. intracranial hemorrhagic stroke / TCH. [Figure 8C] Peptide volcano plot shows (as dark dots) C. dysregulated serum peptides that were significant in healthy controls vs. intracranial hemorrhagic stroke / ICH. [Figure 9]STRING analysis of a unique subset of protein identifiers significantly dysregulated in acute ischemic stroke / AIS compared to intracranial hemorrhagic stroke (ICH). Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) analysis to determine the most common interacting pathways between selected protein identifiers (adjusted P-value (adj.pval) < 0.05, fold change > 2). (A) STRING nodes characteristic of intracranial hemorrhagic stroke (ICH) sera, (B) Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) nodes characteristic of acute ischemic stroke (AIS) sera determined from quantitative comparison. [Figure 10] STRING analysis of a unique subset of protein identifiers significantly dysregulated in acute ischemic stroke / AIS compared to intracranial hemorrhagic stroke (ICH). Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) analysis to determine the most common interacting pathways between selected protein identifiers (adjusted P-value (adj.pval) < 0.05, fold change > 2). (A) STRING nodes characteristic of intracranial hemorrhagic stroke (ICH) sera, (B) Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) nodes characteristic of acute ischemic stroke (AIS) sera determined from quantitative comparison. [Figure 11] Detailed Nature of the Invention and Research Description - Presenting the approach that led to the provision of the present invention. [Example]
[0041] Example 1 This method is novel and simple, with only three steps in sample preparation (Figure 1). It is strongly recommended to prepare all solutions utilizing LC-MS (liquid chromatography with mass spectrometry)-grade water, solvents, reagents, and ultraclean glass and plasticware throughout the method before starting any step of the peptidomics method. Furthermore, all solutions and buffers should be prepared immediately before use and discarded after one week of storage.
[0042] Collection of serum samples Immediately after stroke patients were admitted to the hospital, blood samples were collected by specialized medical staff. Additionally, blood samples from healthy donors served as controls. All study procedures were performed in accordance with the principles of the World Medical Association Declaration of Helsinki. Therefore, the peptidomics study included three different patient sample groups: 1. healthy volunteers (5 donors), 2. patients with acute ischemic stroke (AIS) (5 patients), and 3. patients with intracranial hemorrhagic stroke (ICH) (5 patients). Furthermore, the clotted samples were centrifuged at 3000 × g for 10 minutes at 4°C. The serum (supernatant) was immediately stored at -80°C until analysis.
[0043] Preparation of buffer solutions The first step was to prepare all necessary buffers on the same day as the peptidomic sample preparation.
[0044] First buffer: pH 3.1–3.6 citrate-phosphate buffer (0.131 M citric acid / 0.066 M NaHPO, NaCl 150 mM) adjusted to pH 3.1–3.6 with 1 M NaOH prepared in LC-MS (liquid chromatography with mass spectrometry) grade water.
[0045] A second solution was prepared: 0.2 volume / volume (v / v) % formic acid / methanol (final percentage of formic acid in methanol was 0.2%).
[0046] A third solution was prepared: 0.2 volume / volume (v / v)% formic acid / water (final percentage of formic acid in water was 0.2%).
[0047] A fourth wash buffer was prepared: water / 5% volume / volume (v / v) methanol / 0.2% volume / volume (v / v) formic acid (final percentages of methanol and formic acid in water were 5% and 0.2%, respectively).
[0048] A fifth elution buffer was prepared: water / 80% volume / volume (v / v) methanol / 0.2% volume / volume (v / v) formic acid (final percentages of methanol and formic acid in water were 80% and 0.2%, respectively).
[0049] Peptidome sample preparation One hundred microliters of serum was collected from each sample and incubated with 2.9 ml of citrate-phosphate buffer, pH 3.1-3.6 (pH 3.3 for optimal results), at a temperature below 4-8°C, e.g., on ice, for 3-5 minutes (gentle manual centrifugation is required to dissociate protein-peptide interactions). All of the above mixtures were mixed and centrifuged, e.g., at 5,000 x g and 4°C for 5 minutes. The supernatant was collected and diluted 1:1 with 3 ml of 0.2% volume / volume (v / v) FA in water for LC-MS (liquid chromatography with mass spectrometry).
[0050] One-step peptide purification For peptide purification, all diluted samples (approximately 6 ml each) were subjected to an Oasis cartridge (hydrophilic-lipophilic balance column, 30 mg, Waters), and all detailed steps are as follows:
[0051] i) Cartridge Conditioning: Cartridges were conditioned with 0.2 volume / volume (v / v) % formic acid / methanol (in this solution the final percentage of formic acid in methanol was 0.2%), 1 ml / cartridge.
[0052] ii) Cartridge equilibration: Cartridges were equilibrated with 0.2 volume / volume (v / v) % formic acid / water (in this solution the final percentage of formic acid in water was 0.2%), 1 ml / cartridge.
[0053] iii) Sample loading: In this step 6 ml of sample, which is the supernatant collected after centrifugation, was loaded onto the cartridge (hydrophilic-lipophilic balance column).
[0054] iv) Washing step: The cartridge was washed three times with water containing 5% volume / volume (v / v) methanol / 0.2% volume / volume (v / v) formic acid (in this solution the final proportions of methanol and formic acid in water were 5% and 0.2%, respectively), with a volume of 1 ml per cartridge for each wash.
[0055] v) Elution: Bound material (peptides) was eluted with 1 ml of water containing 80% v / v methanol / 0.2% v / v formic acid (in this solution, the final proportions of methanol and formic acid in water were 80% and 0.2%, respectively), followed by dilution to water / 40% v / v methanol / 0.2% v / v formic acid (in this solution, the final proportions of methanol and formic acid in water were 40% and 0.2%, respectively).
[0056] Peptide preparation for mass spectrometry The serum peptide samples obtained after peptide purification were centrifuged through a 3 kilodalton / kDa molecular weight cutoff (MWCO) ultrafiltration device (Amicon Ultra2 centrifugal filter, Ultracel-3k, 2 mL, Millipore). Briefly, the Amicon-Ultracel-3 kilodalton / kDa centrifugal filter was rinsed before use (the 3 kDa ultrafiltration device was rinsed with 40 volume / volume (v / v)% methanol in LC-MS (liquid chromatography with mass spectrometry)-grade water, 1 mL of water per column, and centrifuged at 4000 × g for 30 min at 4 °C). Next, the 3 kDa ultrafiltration device was inserted into a new collection tube. Finally, the eluted and diluted peptide solution (approximately 2 ml) was pipetted into the 3 kDa ultrafiltration device, taking care not to touch the membrane with the pipette tip. The 3 kDa ultrafiltration device was then centrifuged at 4000 × g for 120–135 min at 4 °C. The filtrate (peptides) was collected and then lyophilized to dryness. All dried peptide samples were stored at -80°C until mass spectrometry analysis.
[0057] To remove high-abundance proteins and simultaneously elute serum peptides, which were further separated by an ultrafiltration approach using a 3 kilodalton (3 kDa) molecular weight cutoff (MWCO) filter (hydrophilic-lipophilic balance), the third step was MS (mass spectrometry) measurement and data analysis. Each patient's serum peptides were measured three times, once by DDA (data-dependent acquisition) and twice by DIA-MS (data-independent acquisition)-MS runs.
[0058] Example 2 LC-MS / MS (liquid chromatography with tandem mass spectrometry) analysis: Eluted and dried serum peptide samples were resuspended in 30 μl of loading buffer consisting of 0.08% trifluoroacetic acid (TFA) in water containing 2.5% acetonitrile (ACN). Indexed retention time (iRT) peptides (Biognosys) were added according to the manufacturer's guidelines. 6 μl of the dissolved sample was then injected into an UltiMate™ 3000 RSLCnano liquid chromatograph (Thermo Scientific) coupled online to an Orbitrap Exploris 480 mass spectrometer (Thermo Scientific). A μ-precolumn C18 trap cartridge (300 μm inner diameter and 5 mm length) packed with C18 PepMap100 sorbent containing PepMap 5 μm sorbent (P / N: 160454, Thermo Fisher Scientific) was used to concentrate and desalt the serum peptides using the loading buffer at a flow rate of 5 μl / min. Peptides were then eluted on a 75 μm internal diameter, 150 mm long fused silica analytical column (P / N: 164534, Thermo Scientific) packed with PepMap 2 μm sorbent. Separation of analytical peptides was achieved by a nonlinear increase in mobile phase B (0.1% formic acid (FA) in ACN) in mobile phase A (0.1% formic acid (FA) in water). The nonlinear gradient started at 2.5% B and increased linearly to 35% B in 80 min, followed by a linear increase to 60% B over the next 15 min at a flow rate of 300 nl / min. Serum peptides eluting from the column were ionized using a nanoelectrospray ion source (NSI) and introduced into an Exploris 480 (Thermo Scientific).
[0059] DDA (Data-Dependent Acquisition) Analysis A data-dependent acquisition method was developed in our laboratory. Briefly, data were acquired on an Exploris 480 in data-dependent peptide mode. Full scans were run in profile mode at 120,000 resolution, scanning the precursor range from m / z 350Th to m / z 1650Th. The normalized AGC (automatic gain control) target was set to 300% with a maximum injection time of 100 ms. After each full scan, the top 20 most intense precursor ions were fragmented and their MS / MS spectra were acquired. Dynamic mass exclusion was set 20 s after the first precursor ion fragmentation. Precursor isotopologues were excluded, and the mass tolerance was set to 10 parts per million (ppm). The minimum precursor ion intensity was set to 3.0e3, and only precursor ion charge states between +1 and +6 were included in the experiment. The precursor isolation window was set to 1.6Th. Normalized collision energy type in fixed collision energy mode was selected. Collision energy was set to 30%. Orbitrap resolution was set to 60000. Normalized AGC (automatic gain control) target was set to 100% with automatic setting of maximum injection time, and data type was centroid.
[0060] DIA (Data Independent Acquisition) Analysis The LC (liquid chromatography) separation parameters were kept identical to the DDA (data-dependent acquisition) acquisition. The Orbitrap Exploris 480 mass spectrometer was operated in positive polarity data-independent mode (DIA) with a full scan in profile mode at 60,000 resolution. The full scan range was set from m / z 350Th to m / z 1450Th, and the normalized AGC (automatic gain control) target was set to 300% with a maximum injection time of 100 ms. Each DIA (data-independent acquisition) cycle was performed with the acquisition of 62 precursor windows / scan events. The DIA (data-independent acquisition) precursor range was set from m / z 350Th to m / z 1100Th with a window width of 12Th and a window overlap of 1Th. The normalized collision energy type in fixed collision energy mode was selected to fragment precursors contained within each isolation window. Collision energy was set to 30% and orb trap resolution in DIA (data independent acquisition) mode was set to 30000. Normalization AGC (automatic gain control) target was set to 1000% with automatic setting of maximum injection time and data type was profile.
[0061] First, to illustrate the relevance of having a mass-appropriate serum peptide isolation protocol and a mass spectrometry-based peptidomics research method, we compared DIA (data-independent acquisition) and DDA (data-dependent acquisition) based on qualitative results (Venn diagram of DIA vs. DDA, Figure 2). The Venn diagram confirms the advantage of using any method while combining all patient identities (IDs) between both approaches. While the majority of identities (IDs) overlap, there are also unique hits in both DIA (data-independent acquisition) and DDA (data-dependent acquisition).
[0062] Example 3 Data analysis First, we applied a multi-search engine strategy to comprehensive qualitative analysis of serum peptidomes. Data-dependent acquisition (DDA) data were centroided and converted to mzML and mzXML formats using MSconvert. The converted MS data were searched against the Homo sapiens SwissProt+TrEMBL reference database, concatenated with indexed retention time (iRT) peptide sequences (Biognosys), decoy inverted target sequences, and contaminants, using MSFragger 3.4, integrated into the Fragpipe suite, and Comet, integrated into Trans-Proteomic Pipeline (TPP) 5.2.0. The search database was constructed in the Fragpipe (v.15) suite. The following search settings were used for MSFragger: precursor mass tolerance was set to + / - 8 parts per million (ppm), fragment mass tolerance was set to 10 parts per million (ppm), enzyme digestion was set to nonspecific, and peptide length was set to 7–45 amino acids. The peptide mass range was set to 200–5000 daltons. Variable modifications were set to methionine oxidation, protein N-terminal acetylation, and cysteinylation of cysteines. The data were mass-recalibrated, and the fragment mass tolerance was adjusted using the automatic parameter optimization settings. The output file format was set to pep.XML. The following search settings were used for Comet: precursor mass tolerance was set to + / - 8 parts per million (ppm), and fragment mass tolerance was set to 8 parts per million (ppm). The remaining settings were identical to those for MSFragger. The search results (pep.XML files) were processed with PeptideProphet and iProphet as part of the Trans-Proteomic Pipeline (TPP) 5.2.0. The resulting recalculated pep.XML files, including peptide iprobabilities (iPROB values), were further processed in Skyline-daily (64-bit, 20.1.9.234).
[0063] The spectral library was constructed from the recalculated pep.XML file using Skyline-daily (64-bit, 20.1.9.234). In "Peptide settings", the "Library" tab was selected. The "Build" option was chosen, and in the newly popped-up window, the "Cut-off score" was set to 0.99 (considering only peptides with a false discovery rate (FDR) < 0.01 or an iPROB value > 0.99), and "Biognosys-11 (iRT-C18)" was selected as the indexed retention time standard peptide. All the recalculated pep.XML files were loaded, and the process of constructing the spectral library was started. Next, in the "Digestion" tab of the "Background proteome" section, the "add" option was selected. In the newly popped-up window, the background proteome was created from the FASTA file (indexed retention time peptide sequences (Biognosys), reverse target sequences, and the SwissProt+TrEMBL reference protein database concatenated with contaminants) that was previously used as a reference search library. Then, in the "Libraries" section of the "Library" tab, the newly constructed spectral library was selected, and the "Explore" button was clicked to explore the library. In the library window, the "Associate proteins" option was checked, and the "Add all" button was clicked. All peptides including non-unique peptides were added to this document. Skyline automatically generated a retention time calculator based on the indexed retention time peptide retention time values observed in the recalculated pep.XML file. The Skyline file was saved to the PC hard disk drive together with the newly constructed spectral library (.blib format).
[0064] Each patient serum sample (15 patients total) was measured using an optimized DDA (data-dependent acquisition) / DIA / SWATH (sequential window acquisition of all theoretical mass spectra) method. Each serum sample was run in one DDA (data-dependent acquisition) and two SWATH (sequential window acquisition of all theoretical mass spectra) / DIA (data-independent acquisition) technical replicates to ensure the highest LC-MS / MS (liquid chromatography with tandem mass spectrometry) assay quality. The DDA (data-dependent acquisition) data were used to generate spectral libraries used for qualitative serum peptidomics as well as for later extraction of quantitative data from DIA / SWATH (sequential window acquisition of all theoretical mass spectra) files. The bar graph diagram (Figure 3) of identified peptides across (5 DDAs (data-dependent acquisitions) vs. 5 DIAs / group) represents the entire serum peptidome analyzed in the three patient groups using our multi-search engine approach described in the Methods section. Error bars represent the variation in serum peptide counts between patients (per patient, not per measurement run). The gray shading (Figure 3) represents the number of identified peptides in the two data types, i.e., DDAs (data-dependent acquisitions) and DIAs (data-independent acquisitions), searched by both Comet and MSfragger, respectively. The Comet search engine performs the best search DIA (data-independent acquisition) data, while MSFragger performs the full DDA (data-dependent acquisition) data search. Search engine performance is highly dependent on search settings; for comparison purposes, both search engines should be set to optimal but equal settings.
[0065] The number of peptides identified in stroke patients was significantly lower than that in healthy controls (Figure 3). In further analysis, we considered only serum peptides quantified in patient and control samples. The reproducibility of the identifications was demonstrated (Figure 3). Furthermore, the relatively high standard error (STDEV > 17.3) in the identification of multiple peptides within a patient group likely corresponds to heterogeneity between patients, as further evidenced by the Venn diagram plot, which showed overlap in serum peptide identities (IDs) within patient groups but variability between patient groups (Figure 4). This observation may be present in any dataset and will affect the results, but we will later demonstrate that our approach can address this issue.
[0066] Example 4 Another aspect of the present invention provides comprehensive identification of serum peptides from patient samples. In our data-dependent acquisition (DDA) analysis (Figure 4), we identified a total of 15,000-17,000 serum peptides with reliable statistical filtration (iProphet iprob values >0.99% in 15 acute ischemic stroke / AIS, intracranial hemorrhagic stroke / ICH, and control / CON samples).
[0067] Figure 4 shows that a total of approximately 7000, 5800, and 4200 serum peptide sequences were identified in control / CON, acute ischemic stroke / AIS, and intracranial hemorrhagic stroke / ICH samples, respectively, with approximately 4000 (90%) commonly identified exfoliated peptides. The potential of our novel approach for use with large clinical sets is demonstrated by the high degree of overlap of all peptides from the three experimental groups (Figure 4).
[0068] Example 5 This study highlights that other important features of serum peptidomic analysis in our data-dependent acquisition (DDA) screening are peptide length distribution and molecular weight range (Figure 5). Peptide length distribution depends on the target of interest, and short peptides (<3 kDa) and long peptides (<10 and <30 kDa) can be easily tuned and enriched by using ultrafiltration membranes with cutoffs of 3, 10, and 30 kDa, respectively.
[0069] In this analysis, a 3 kilodalton (kDa) filter was used during the peptide purification step, resulting in a distribution of 7 to 30 amino acid residues with a median of 15 amino acid residues (Figure 5). Therefore, the molecular weight range of the enriched serum peptides indicated a mass distribution in the range of 1 to 3 kilodaltons (kDa).
[0070] Example 6 A further application is based entirely on serum peptide quantification. In quantification experiments, we first used data-dependent acquisition (DDA) data to create an appropriate spectral library without losing many important serum peptides and simultaneously introducing too many low-confidence serum peptides. Then, we used only data-independent acquisition (DIA) data to extract quantitative information about the serum peptides included in the spectral library. Furthermore, we consider DIA data to be the only option for adequately quantifying serum peptides due to its fully reproducible multiplexing. Furthermore, as previously suggested, DIA data can be used for both identification and precise quantification of serum peptides. To accurately quantify stroke-specific markers, we developed several methods with distinctly different window widths. Peptide precursor ion distributions were first screened from serum samples using an optimized DDA method. The present invention determined the most densely populated portion of the precursor range, encompassing the majority of observed peptide masses. The remaining DIA (data-independent acquisition) method parameters, such as normalization AGC (automatic gain control) target, injection time, and window overlap, were optimized to keep cycle times to a maximum of 3.5 seconds. Such a method cycle, combined with the described chromatographic conditions (Methods section), yields a peptide peak with at least 10 data points, the minimum recommended by the FDA (U.S. Food and Drug Administration), while maintaining sufficient selectivity and sensitivity of the method.
[0071] Quantitative peptidomics data extraction Continuing with the Skyline file created in the previous step is highly recommended to ensure proper extraction settings. Quantitative values reflecting peptide abundance (peptide peak area) in the sample were extracted from the DIA (Data Independent Acquisition) data in Skyline-daily (64-bit, 20.1.9.234) based on the peptide-product ion pairs (transitions) listed in the spectral library from the previous step. In the "Peptide settings" tab, "Max. missed cleavages" was set to 0. Next, in the "Filter" subtab, the maximum peptide length was set to 4–200 amino acids. The "Exclude N-terminal AAs" option was set to 0. The "Auto-select all matching peptides" function was selected. In the "Modifications" subtab, no modifications were selected. In the "Library" subtab, the spectral library created in the previous step was selected in the "Libraries" window. Other settings within the tab were left at their defaults. In the "Transition settings" tab, the "Prediction" subtab was left at its default settings. In the "Filter" subtab, peptide precursor charges of +1, +2, +3, +5, and +6 were specified. The ion charge was set to +1 and +2. The ion type was set to y and b. Product ion selection was set as follows: "ion 4" was selected in the "From:" window and "last ion" was selected in the "To:" window. The "Auto select all matching transition" option was selected at the bottom of the subtab, and the "N-terminal to Proline" option was selected in the "Special ions:" menu. In the "Library" subtab, the "Ion match tolerance" window was set to 0.05 m / z, and only peptides with at least four product ions were retained for analysis. Furthermore, if more product ions per peptide were available, the six most intense ones were selected from the filtered product ions.The "From filtered ion charges and types" function was selected. The "Instrument" subtab included product ions between m / z 350 and m / z 1100. The "Method match tolerance m / z" window was set to a mass tolerance of 0.055 m / z. In the full scan subtab, mass spectrometry (MS) filtering was set to "none." In the tandem mass spectrometry (MS / MS) filtering subsection, the "Data-Independent Acquisition (DIA)" option was selected in the "Acquisition method" window, and the "Orbitrap" option was selected in the "Product mass analyzer" section. The "Add" option was selected in the "Isolation scheme:" section, and the "Prespecified isolation windows" option was selected in the "Edit isolation scheme" pop-up window. The "Import" isolation window was then selected from the DIA raw file options, and the isolation scheme was read from the DIA raw file. Returning to the "Full-Scan" subtab in the "Resolving power:" section, we set the resolution to 30,000 at 200 m / z. In the retention time filtering subsection, we specified "Use only scans within 5 minutes of tandem mass spectrometry (MS / MS) identifications (IDs)." In the "Document" section of the "Advanced" window under the "Refine" subtab, we set the "Min transitions per precursor" option to 4. Empty proteins were removed from the document. We added an equal number of reverse sequence decoys via the "Add Decoys" function in the "Refine" subtab. In the "Add Decoy Peptides" pop-up window, we selected the reverse sequence decoy generation method to create decoy peptides.The DIA (data-independent acquisition) ".raw" files were then imported. An mProphet model for reintegrating peptide product ion peak boundaries in the product ion chromatogram was trained as follows: In the "Refine" subtab, the "Reintegrate" function was selected. In the "Reintegrate" popup window for "Peak scoring model:," the "Add" option was selected. The mProphet model was then launched in "Edit Peak Scoring Model." The mProphet model was trained using the target and decoys. Score rows with negative "Weight" and / or "Percentage Contribution" (highlighted in red) were deleted, and the mProphet model was then retrained. A new mProphet peak scoring model was then selected in the "Edit peak scoring Model" window and applied to the new peak boundary reintegration. A quantitative report for downstream analysis was generated using the "Export report" function, which exports a summary of all required dependencies to MSstats 4.0.1. Data-Independent Acquisition (DIA) Statistical Data Analysis in the MSstats R Module. Statistical analysis of Skyline-extracted quantitative data-independent acquisition (DIA) data was performed using the R (version 4.0.0) package MSstats 4.0.1. Protein columns were combined with peptide sequence columns to preserve analysis at the peptide level; this formatting prevented the summation of peptide intensities from a single protein in MSstats. Extracted peaks were reduced by filtering with a q-value <0.01 cutoff in mProphet. The "SkylinetoMSstatsFormat" function was configured to retain proteins with a single feature and convert Skyline output to MSstats-compatible input. Furthermore, peptide intensities were log2-transformed and quantile-normalized.Differential serum peptide quantification across intracranial hemorrhagic stroke (ICH) and acute ischemic stroke (AIS) conditions was performed pairwise via a mixed-effects model implemented in the "groupComparison" function in MSstats. p-values were adjusted using the Benjamini-Hochberg method, and the resulting matrix was exported for downstream analysis. The log2-transformed, quantile-normalized data matrix was exported to ProBatch 1.8.0 running under R (version 4.0.0) to generate sample correlation heatmaps, dendrograms, and analyses. Heatmaps were created with the Heatplus 3.0.0 package, volcano plots with the eulerr 6.1.1 package, and bar graphs with plyr 1.8.6 and ggplot2 3.3.5, all in R (version 4.0.0).
[0072] Using our novel approach, approximately 1900-2100 serum peptides (extraction methods are described in the "Methods" section) were successfully quantified by 30 LC-MS / MS (liquid chromatography with tandem mass spectrometry) analyses of 15 clinical samples with data-independent acquisition (DIA). Several peptides were quantified across all patient groups (acute ischemic stroke / AIS, intracranial hemorrhagic stroke / ICH, and control / CON), showing a wide range of fold changes (2-135 fold change / FCH) across the compared patient groups. Notably, the signal intensity of test peptides in patient groups was at least 2-fold higher (≥2) or at least 2-fold lower (≤-2) than that of control samples. Our results show peptides that are completely absent in one of the conditions. This phenomenon provides infinite variation and a p-value of 0. Comparison with matched control samples allowed identification (qualitative analysis) of at least 15,000 serum peptides and quantification (quantitative analysis) of at least 1,900 serum peptides. Subsequently, our quantitative serum peptidomics pipeline acquired and extracted quantitative information about serum peptides from the DIA (data-independent acquisition) data. A correlation heat map of MS (mass spectrometry) runs (Figure 6) was then plotted to compare and correlate 30 individual DIA (data-independent acquisition)-MS (mass spectrometry) runs of serum peptides. As expected, we observed the greatest correlation between technical replicates (diagonal squares), demonstrating the excellent performance of the LC-MS / MS (liquid chromatography with tandem mass spectrometry) system (R>0.9). Furthermore, the correlation heat map of MS (mass spectrometry) runs suggests that the serum peptidome of healthy donors is more correlated than that of stroke patients (top left square). However, there are clear stratified differences between serum peptides of stroke patients and healthy donors. This identification demonstrated that our novel approach achieves significant results for the peptidomics platform. The lower right corner of the heatmap in Figure 6 suggests a relatively good correlation (R of approximately 0.7) within the intracranial hemorrhagic stroke (ICH) patient group. On the other hand, acute ischemic stroke (AIS) patients show the highest peptidome heterogeneity.The sample correlation heatmap suggests only minor serum peptidome differences between ICH and AIS, but at the same time, our peptidomic platform reveals even minute differences in serum. Overall, the relatively low correlation coefficients between study subjects, primarily in AIS, may correspond to patient heterogeneity.
[0073] Next, quantitative serum peptidomic signatures were correlated with patient groups. Finally, unsupervised hierarchical clustering of 30 DIA (data-independent acquisition) LC-MS / MS (liquid chromatography with tandem mass spectrometry) runs was performed. The presented peptide heatmap and unsupervised hierarchical clustering (Figure 7) provide an overview of quantitative peptidomic data analysis of serum peptidomic DIA-MS runs (data-independent acquisition-mass spectrometry) between stroke patients and healthy donors.
[0074] Strikingly, similar log2 peptide intensity patterns were observed among the compared groups, resulting in a perfect clustering of the serum peptidomes of healthy donors and stroke (AIS, ICH). Interestingly, the hierarchical clustering function also reliably distinguishes the serum peptidome of acute ischemic stroke (AIS) from that of intracranial hemorrhagic stroke (ICH). The serum peptidomes of the two outliers, which clustered differently, may result from the aforementioned interpatient heterogeneity or different patient clinical histories, which were not explored. These data are in excellent agreement with the sample correlation heatmap (Figure 6).
[0075] Example 7 Next, the present invention proceeded to provide peptide quantification between the compared patient groups. Significantly dysregulated peptides (adj.pval) ≦0.05, fold change ≧2) were visualized as a volcano plot, as shown in FIG.
[0076] The Volcano plot suggests that peptide quantification provides a list of peptides that significantly stratify between the compared patient groups, and therefore it was successful in screening for statistically significant serum peptides.
[0077] Example 8 To further understand the biological significance inferred from the serum peptidome quantitative data of stroke patients, we performed a multi-bioinformatics approach, including Gene Ontology (GO) enrichment analysis (molecular function, biological process, and component analysis), Search Tool for the Retrieval of Interacting Genes / Proteins (STRING), and keyword enrichment (Uniport). First, we created an interactome map using Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) to determine characteristic nodes enriched in the respective comparisons of intracranial hemorrhagic stroke / ICH and acute ischemic stroke / AIS. Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) aims to specifically discriminate between strokes (intracranial hemorrhagic stroke / ICH and acute ischemic stroke / AIS). STRING analysis identified significantly dysregulated protein identifiers in acute ischemic stroke / AIS and intracranial hemorrhagic stroke (ICH). The filtered subset of unique protein identifiers corresponding to significantly upregulated peptides (adjusted P value (adj.pval) ≤ 0.05, fold change ≥ 2) from the comparison of acute ischemic stroke / AIS versus intracranial hemorrhagic stroke / ICH, and from the comparison of intracranial hemorrhagic stroke / ICH versus acute ischemic stroke / AIS were converted to Uniport protein identifiers. Identifiers common to both the upregulated and downregulated interaction maps were removed, and the unique identifiers were subjected to Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) analysis. The analysis revealed interesting Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) nodes characteristic of intracranial hemorrhagic stroke / ICH serum. Figure 9 (A) and (B) for acute ischemic stroke (AIS) serum were determined from their quantitative comparison.
[0078] This is the type of biological validation further performed by literature searches. More than 2,000 quantified peptides were mapped to 200 annotated protein precursors. As shown in Figures 9 and 10, STRING analysis revealed that key terms (references to exact pathways, if present) related to extracellular exosomes, blood microparticles, platelet alpha granule lumen, and extracellular space were most significantly enriched. Therefore, the Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) interaction network provides another powerful method for identifying key enriched protein nodes that may determine the predictive role of peptides. Thus, our novel approach for serum peptidomics sample preparation is comprehensive, qualitative, and quantitative. We primarily considered serum samples from different stroke (acute ischemic stroke / AIS), intracranial hemorrhagic stroke / ICH), and healthy volunteers. Therefore, it revolutionizes the field of early diagnosis and presents a roadmap for the development of point-of-care diagnostics (POCD). Ultimately, this will improve the quality of life and be applicable worldwide. Therefore, this invention will bring a breakthrough to the point-of-care diagnostic (POCD) platform and revolutionize the medical field. It can be immediately used for all types of (human / non-human) control / disease serum samples in the future.
[0079] References [1] R. Richter et al., “Composition of the peptide fraction in human blood plasma: database of circulating human peptides,” J. Chromatogr. B. Biomed. Sci. App., vol. 726, no. 1-2, pp. 25-35, Apr. 1999, doi: 10.1016 / s0378-4347(99)00012-2. [2] J. Yang et al., “Serum peptidome profiling in patients with lung cancer,” Anat. Rec. Hoboken NJ 2007, vol. 293, no. 12, pp. 2027-2033, Dec. 2010, doi: 10.1002 / ar.21267. [3] Y. Araki et al., “Clinical peptidomic analysis by a one-step direct transfer technology: its potential utility for monitoring of pathophysiological status in female reproductive system disorders,” J. Obstet. Gynaecol. Res., vol. 39, no. 10, pp. 1440-1448, Oct. 2013, doi: 10.1111 / jog.12140. [4] L. Yin, Y. Huai, C. Zhao, H. Ding, T. Jiang, and Z. Shi, “Early Second-Trimester Peptidomic Identification of Serum Peptides for Potential Prediction of Gestational Diabetes Mellitus,” Cell. Physiol. Biochem., vol. 51, no. 3, pp. 1264-1275, 2018, doi: 10.1159 / 000495538. [5] A. A. Abdelati, R. A. Elnemr, N. S. Kandil, F. I. Dwedar, and R. A. Ghazala, “Serum Peptidomic Profile as a Novel Biomarker for Rheumatoid Arthritis,” Int. J. Rheumatol., vol. 2020, p. e6069484, Aug. 2020, doi: 10.1155 / 2020 / 6069484. [6] Z. Miao, K. Ding, S. Jin, L. Dai, C. Dai, and X. Li, “Using serum peptidomics to discovery the diagnostic marker for different stage of ulcerative colitis,” J. Pharm. Biomed. Anal., vol. 193, p. 113725, Jan. 2021, doi: 10.1016 / j.jpba.2020.113725. [7] H. Ay, “Biomarkers for stroke diagnosis,” WO2014121252A1, Aug. 07, 2014 Accessed: Dec. 17, 2022. [Online]. Available: https: / / patents.google.com / patent / WO2014121252A1 / en [8] T. Garcia-Berrocoso and J. Montaner Vilallonga, “Biomarkers for the Prognosis of Ischemic Stroke,” May 01, 2014 Accessed: Dec. 17, 2022. [Online]. Available: https: / / patentscope.wipo.int / search / en / detail.jsf?docId=WO2014064202 [9] C. Watson, M. Ledwidge, K. Mcdonald, and J. Baugh, “Biomarkers of cardiovascular disease including lrg,” WO2011092219A1, Aug. 04, 2011 Accessed: Dec. 17, 2022. [Online]. Available: https: / / patents.google.com / patent / WO2011092219A1 / en?oq=PCT%2fEP2011%2f051088
[10] P. Broberg, T. Fehniger, C. Lindberg, and G. MARKO-VARGA, “Peptides as biomarkers of copd,” WO2006118522A1, Nov. 09, 2006 Accessed: Dec. 17, 2022. [Online]. Available: https: / / patents.google.com / patent / WO2006118522A1 / en?oq=PCT%2fSE2006%2f000507
[11] J. Klein, M. L. Merchant, R. Ouseph, and R. A. Ward, “Peptide biomarkers of cardiovascular disease,” US20140243432A1, Aug. 28, 2014 Accessed: Dec. 17, 2022. [Online]. Available: https: / / patents.google.com / patent / US20140243432A1 / en?oq=US+2014%2f0243432+A1
[12] K. I. VADAKKAN, “Biomarkers for rapid detection of an occurrence of a stroke event,” WO2014110674A1, Jul. 24, 2014 Accessed: Dec. 17, 2022. [Online]. Available: https: / / patents.google.com / patent / WO2014110674A1 / en?oq=PCT%2fCA2014%2f050023
[13] J. M. Villalonga, A. S. ORIOL, and L. R. PASCUAL, “Biomarkers for stroke prognosis,” WO2021009287A1, Jan. 21, 2021 Accessed: Dec. 17, 2022. [Online]. Available: https: / / patents.google.com / patent / WO2021009287A1 / en?oq=PCT%2fEP2020%2f070152
[14] T. Linke, S. Doraiswamy, and E. H. Harrison, “Rat plasma proteomics: Effects of abundant protein depletion on proteomic analysis,” J. Chromatogr. B, vol. 849, no. 1, pp. 273-281, Apr. 2007, doi: 10.1016 / j.jchromb.2006.11.051.
[15] “Albumin depletion of human plasma also removes low abundance proteins including the cytokines - Granger - 2005 - PROTEOMICS - Wiley Online Library.” https: / / analyticalsciencejournals.onlinelibrary.wiley.com / doi / 10.1002 / pmic.200401331 (accessed Dec. 17, 2022).
Claims
1. 1. A method for sample preparation for the qualitative and quantitative analysis of the peptidome in serum based on amino acid sequencing using a tandem mass spectrometry (MS) tool with signaling intensity measurements of peptides, said method comprising the following steps: a) a step of concentrating peptides in the serum sample to be analyzed by incubating the sample with a citrate-phosphate buffer solution of pH 3.1-3.6 in water of purity 99.9% or higher, the volume ratio of the buffer to the serum being 29:1, and mixing the buffer and the serum at a temperature in the range of 4-8°C to dissociate peptides from high-concentration serum proteins; b) removing proteins of 15 kDa or larger from the concentrated peptides using a hydrophilic-lipophilic balance column and simultaneously purifying the peptides to obtain peptides of less than 15 kDa; c) washing the hydrophilic-lipophilic balance column with high-grade water of 99.9% purity or higher together with 0.2% formic acid (v / v) to remove unbound and low-bound peptides from the column; d) eluting the serum peptides bound to the hydrophilic-lipophilic balance column by using water / 80% methanol / 0.2% formic acid (v / v); e) filtering the sample through a 3 kDa molecular weight filter to collect peptides less than 3 kDa; f) providing lyophilization of the obtained peptides by drying the eluted peptides by lyophilization for tandem mass spectrometry; A method comprising:
2. 2. The method of claim 1, wherein step f) is followed by the following: sequencing the isolated peptides based on their amino acid sequence by tandem mass spectrometry, followed by performing intensity signaling measurements on the tested serum samples compared to matched control serum samples.
3. 3. The method of claim 2, wherein the identification (qualitative analysis) of at least 15,000 serum peptides and the quantification (quantitative analysis) of at least 1,900 serum peptides can be performed based on the comparison with matched control samples.
4. The method according to claims 1 to 3, wherein the peptides finally obtained for the analysis may have different sizes according to the molecular weight filter used to collect the peptides. Therefore, the permeability of the filter is not a limitation of the method.
5. The following steps are included in the Skyline and MSStats pipelines: a) searching MS / MS DDA (data-dependent acquisition) data with two or more search engine strategies to generate a comprehensive spectral library; b) extracting quantitative peptidomic data (independent MS / MS data, DIA) using said comprehensive spectral library; c) extracting a chromatogram of product ions and integrating the area of the product ion peaks (area under the curve, AUC); d) controlling the false discovery rate (FDR) of the quantitative extraction / analysis by determining the FDR values of the peaks of the mProphet module; e) Inferring product ion peak areas for peptides and performing their sum peak area statistics in the MSstats module with a peptide-centric view; f) further downstream analysis, e.g., graphical data visualization including volcano plots, dendrograms, heat maps, functional stroke serum peptide enrichment maps using quantitative matrices from MSstats; A data analysis workflow for quantitative analysis of tandem mass spectrometry sequenced peptides obtained according to claims 1 to 3, comprising: said downstream analysis is therefore not a limitation of said method.
6. 6. The method of claim 5, wherein quantitative analysis of the peptides sequenced by tandem mass spectrometry can be performed if the signal intensity of the tested peptide is at least two-fold higher (≧2) or at least two-fold lower (≦−2) than the signal intensity of a matched control sample.