Methods and systems for high-throughput proteomics discovery

WO2026170078A1PCT designated stage Publication Date: 2026-08-13SEER INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-06
Publication Date
2026-08-13

Smart Images

  • Figure US2026014400_13082026_PF_FP_ABST
    Figure US2026014400_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Described herein are data independent analysis search engines and high-throughput data analysis pipelines for identifying protein structure matches in mass spectrometry data acquired by data independent acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

WSGR Docket No. 53344-800.601METHODS AND SYSTEMS FOR HIGH-THROUGHPUT PROTEOMICS DISCOVERYCROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 755,210, filed February 6, 2025; U.S. Provisional Application No. 63 / 817,022, filed June 3, 2025; and U.S.Provisional Application No. 63 / 913,868, filed November 7, 2025; each of which is incorporated herein by reference in its entirety.BACKGROUND

[0002] Proteomics is a central tool to understanding cellular processes and disease mechanisms. Data-independent Acquisition (DIA) is a discovery mass spectrometry technique that allows for sensitive and reproducible protein quantification across complex biological samples. However, the performance and scalability of DIA relies heavily on the performance of the search engines used to analyze the massive datasets generated. Whereas affinity proteomic studies can already scale to hundreds of thousands of samples, state-of-the-art DIA proteomics studies rarely exceed a few hundred acquisitions. Recent advances in instrumentation and sample preparations now make large-scale DIA measurements feasible, though analysis of single-cell atlases, population proteomics, and clinical cohorts is increasingly limited by computational throughput. Bridging this gap requires search engines and pipelines designed explicitly for population-scale mass-spectrometry (MS) data. Further, while large-scale mass spectrometry (MS) experiments enable comprehensive identification and quantification of proteins in complex biological samples, datasets with thousands of injections present challenges in maintaining sensitivity, quantitative accuracy, and a well-controlled false discovery rate (FDR). Traditional pipelines often struggle to scale while preserving statistical rigor and confidence in results. Although these tools have been foundational in advancing proteomics research, their tightly coupled framework limit their ability to scale and integrate with modem distributed-data processing frameworks. Existing approaches still require global processing on a single node, restricting scalability. Because global processing still occurs on a single node, and CPU, memory and storage requirements grow super-linearly with dataset size, processing experiments larger than a few thousand MS acquisitions becomes impractical. Resource limitations hinder the analysis of large datasets and the application of innovative data science techniques.

[0003] Accordingly, improved data processing methods and approaches are needed.SUMMARY

[0004] According to an aspect of the invention, a method is provided comprising receiving a plurality of retention times from a mass spectrometer, processing the plurality of retention times using a model configured to calibrate a retention time of the plurality of retention times, determining a first protein structure match based on the processed plurality of retention times, and training a machine learning model to output a second (or optimized) protein structure match, wherein the machine learning modelWSGR Docket No. 53344-800.601is trained on a dataset comprising the first protein structure match and the plurality of processed retention times. In some embodiments, receiving the plurality of retention times comprises receiving a data file generated by the mass spectrometer (e.g., a data file corresponding to a single mass spectrometer injection). In some embodiments, the mass spectrometer is configured for data independent acquisition (DIA). In some embodiments, the received plurality of retention times are stored in a database (e.g., located in a memory of a computational device). In one embodiment, the database has a R-tree structure (e.g., an original R-tree structure, a packed R-tree structure, or a similar database structure). In some embodiments, the first protein structure match comprises an identified peptide. In one embodiment, the first protein structure match further comprises a corresponding false discovery rate (FDR). In some embodiments, the second protein structure match comprises the first protein structure match and an optimized FDR. In some embodiments, the model configured to calibrate the retention time comprises a retention time (RT) alignment model and an RT window filter. In one embodiment, the RT window filter is derived from the RT alignment model. In some embodiments, the model configured to calibrate the retention time comprises a regression model. In some embodiments, the model configured to calibrate the retention time comprises a linear discriminant analysis (LDA). In some embodiments, the model configured to calibrate the retention time is trained on a dataset comprising empirical retention data and predicted retention data. In one embodiment, the predicted retention data is obtained from a target library. In one embodiment, the targets in the target library are peptides. In some embodiments, receiving the plurality of retention times further comprises receiving a plurality of mass-to-charge (m / z) values, wherein each retention time of the plurality of retention times is paired with a corresponding m / z value of the plurality of m / z values. Alternatively or additionally, the method may further comprise computing a mass calibration curve and recalibrating the received plurality of corresponding m / z values using the computed mass calibration curve. In one embodiment, computing the mass calibration curve and recalibrating the received plurality of corresponding m / z values is performed after calibrating the plurality of retention times. In some embodiments, determining the first protein structure match is based on the processed plurality of retention times and the recalibrated plurality of corresponding m / z values. In some embodiments, receiving the plurality of retention times further comprises receiving a plurality of signal intensities, wherein each retention time of the plurality of retention times is paired with a corresponding signal intensity of the plurality of signal intensities. In one embodiment, determining the first protein structure match is based on the processed plurality of retention times, the plurality of corresponding signal intensities, and, optionally, the recalibrated plurality of corresponding m / z values. In some embodiments, receiving the plurality of retention times further comprises receiving a plurality of ion mobility values, wherein each retention time of the plurality of retention times is paired with a corresponding ion mobility value. Alternatively or additionally, the method may further comprise computing an ion mobility calibration curve and recalibrating the received plurality ofWSGR Docket No. 53344-800.601corresponding ion mobility values using the computed ion mobility calibration curve. In one embodiment, computing the ion mobility calibration curve and recalibrating the received plurality of corresponding ion mobility values is performed after calibrating the plurality of retention times. In some embodiments, determining the first protein structure match is based on the processed plurality of retention times, the recalibrated plurality of corresponding ion mobility values, and, optionally, the recalibrated plurality of corresponding m / z values and / or the plurality of corresponding signal intensities. In some embodiments, the machine learning model comprises a neural network. In some embodiments, the second protein structure match output by the machine learning model comprises an optimized FDR. Alternatively or additionally, the method may further comprise determining a plurality of first protein structure matches, wherein the machine learning model is trained on the plurality of first protein structure matches, the plurality of processed retention times, and, optionally, the recalibrated plurality of corresponding m / z values and / or the plurality of corresponding signal intensities and / or the recalibrated plurality of corresponding ion mobility values. In one embodiment, each protein structure match of the plurality of first protein structure matches has an FDR that is below a pre-selected threshold. In one embodiment, the pre-selected FDR threshold for the first protein structure matches is 50% FDR or less (e.g., 45%, 40%, 35%, 30%, 25%, or less). Alternatively or additionally, the method may further comprise outputting a list comprising a plurality of second protein structure matches, wherein each second protein structure match passes a pre-selected FDR threshold. In one embodiment, the pre-selected FDR threshold for the plurality of second protein structure matches is user-selected. In some embodiments, the pre-selected FDR threshold for the plurality of second protein structure matches is an FDR of 1% or less.

[0005] According to another aspect of the invention, a method is provided comprising receiving, at a plurality of processing units, mass spectrometry data from at least one mass spectrometer, processing the mass spectrometry data, wherein the processing comprises calibrating a retention time of the mass-spectrometry data, and determining a plurality of protein structure matches. In some embodiments, the processing comprises performing the method as described above. In some embodiments, the plurality of protein structure matches has a false match rate of less than about 0.5%. In some embodiments, the false match rate at a precursor level is less than about 1%. In some embodiments, the false match rate at a protein group level is less than about 1%. In some embodiments, the plurality of processing units comprises at least two nodes. In one embodiment, the calibrating the retention time comprises calibrating an ion retention time based on an average mass calibration curve.

[0006] According to another aspect of the invention, a method for protein identification and / or quantification by mass spectrometry (MS) experiments is provided, comprising obtaining a plurality of mass spectra utilizing data-independent acquisition (DIA), processing by linear discriminant analysis (LDA) input data representing complex biological samples to build a model that correlates predicted retention times with empirical values, applying a retention time window filter based on theWSGR Docket No. 53344-800.601model generated by LDA, recalibrating mass-to-charge (m / z) values for all ions in the plurality of mass spectra using average mass calibration curves across batches, and determining peptide-spectrum matches (PSMs) based at least in part on the model generated by LDA. In some embodiments, the method further comprises a high-throughput proteomics workflow for processing at least 1000 or more data files (e.g., at least 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10,000, 11,000, 12,000, 13,000, 14,000, 15,000, 20,000, 25,000, or more data files) comprising the plurality of mass spectra within a 24-hour period. In some embodiments, the method processes a population-scale proteomics study with an FDR of less than about 50% (e.g., less than about 45%, 40%, 35%, or 30%). In some embodiments, a false discovery rate of the method is less than about 5% (e.g., less than about 4%, 3%, 2%, 1.5%, or 1%). In some embodiments, the LDA processing is performed individually on each MS data file comprising mass spectra of the plurality of mass spectra. In some embodiments, the LDA processing comprises a pairwise analysis of peptide target spectra and corresponding decoy spectra. In one embodiment, the peptide target spectra are a subset of peptide target spectra from a peptide target library. In one embodiment, the peptide target spectra, the decoy spectra, and / or the peptide target library is / are stored in an accessible memory database having an R-tree structure, for example comprising an original R-tree, a packed R-tree, or a similar database structure. In some embodiments, the peptide target spectra and the corresponding decoy spectra are represented by multi-dimensional feature vectors. In one embodiment, the multi-dimensional feature vectors comprise at least 50 features (e.g., at least 75, 100, 110, 120, 130, 140, 150, or more features). In some embodiments, the average mass calibration curves are generated using data output from the LDA processing. In some embodiments, determining the PSMs comprises performing regression analysis. In some embodiments, determining the PSMs comprises performing LDA. In some embodiments, determining the PSMs comprises generating feature vectors representative of each peptide target spectrum of a peptide target library, and scoring the feature vectors against mass spectra of the plurality of mass spectra having recalibrated m / z values. In one embodiment, each feature vector comprises at least 100 features (e.g., at least 120, 130, 140, 150 features; or about 160 to about 200 features, about 170 to about 190 features, or about 180 features). In some embodiments, determining the PSMs is further based at least in part on the retention time window filter and / or the recalibrated m / z values. Alternatively or additionally, the method may further comprise normalization of ion counts or intensities prior to the step of determining PSMs. Alternatively or additionally, the method may further comprise calculating precursor quantities for the determined PSMs. In one embodiment, calculating a precursor quantity comprises summing detected fragment intensities. In one embodiment, summing the detected fragment intensities comprises summing at least the detected fragment intensities over the highest single scan. In some embodiments, summing the detected fragment intensities comprises summing the detected fragment intensities over all scans.WSGR Docket No. 53344-800.601

[0007] According to another aspect of the invention, a method for searching mass spectrometry data is provided, the method comprising obtaining a mass spectrometry (MS) data file comprising retention times and corresponding mass-to-charge (m / z) values, performing a first search of the MS data file using a first plurality of peptide target spectra from a library of peptide targets, thereby identifying a first set of peptide-spectrum matches (PSMs), calculating an optimized retention time (RT) window filter using predicted retention times for target peptides represented in the first set of PSMs and their corresponding observed retention times in the MS data file, computing a mass calibration curve using predicted m / z values for the target peptides in the first set of PSMs and corresponding observed m / z values in the MS data file, recalibrating the m / z values in the MS data file using the computed mass calibration curve, thereby generating a recalibrated MS data file, calculating optimized mass tolerances by performing a series of iterative searches on the recalibrated MS data file, using a second plurality of peptide target spectra from the peptide target library, scoring the peptide target library against the recalibrated MS data file using the optimized RT window filter and the optimized mass tolerances, thereby generating a set of optimized PSMs, and calculating optimized false discovery rates (FDRs) for the optimized PSMs in the set of optimized PSMs using a machine learning algorithm. Alternatively or additionally, the method may further comprise outputting a list of optimized PSMs having optimized FDRs below a pre-selected FDR threshold. In one embodiment, the pre-selected FDR threshold is user-selected. In some embodiments, the pre-selected FDR threshold is an FDR of 1% (e.g., 2.0%, 1.5%, 1%, or less). In some embodiments, the list of optimized PSMs comprises a feature vector and related meta-data for each optimized PSM in the list. In one embodiment, the metadata includes a score (e.g., single-value score) for the feature vector, a sequence for the matched peptide, a retention time for the matched peptide, and optionally ion mobility time for the matched peptide. In some embodiments, the MS data file corresponds to a single MS injection. In some embodiments, the MS data file has an mzML or Parquet file format. In some embodiments, the first plurality of peptide target spectra corresponds to about 0.1% to about 25% of the peptide targets in the peptide target library (e.g., up to about 20%, 15%, 10%, 7%, 5%, 4%, 3%, 2%, or 1% of the peptide targets in the peptide target library). In some embodiments, the first plurality of peptide target spectra comprises about 20,000 peptide target spectra (e.g., at least about 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000 or more peptide target spectra). Alternatively or additionally, the method may further comprise obtaining the peptide target library and storing the peptide target library in a memory database having an R-tree structure. In one embodiment, the R-tree structure is an original R-tree, a packed R-tree, or has a similar database structure. In some embodiments, each peptide target spectrum in the first plurality of peptide target spectra is represented as a vector having a plurality of features. In one embodiment, each vector representing a peptide target spectrum comprises at least 50 features (e.g., at least 75, 100, 110, 120, 130, 140, 150, or more features). Alternatively or additionally, the method may further comprise, prior to building an RT alignment model, generating a decoyWSGR Docket No. 53344-800.601spectrum corresponding to each peptide target spectrum in the first plurality of peptide target spectra. In one embodiment, generating a decoy spectrum corresponding to a target spectrum comprises determining a spectrum corresponding to a mutated peptide sequence derived from a peptide sequence corresponding to the peptide target spectrum. In some embodiments, the decoy spectrum is represented as a vector having a plurality of features. In one embodiment, each decoy vector representing a peptide target spectrum comprises at least 50 features (e.g., at least 75, 100, 110, 120, 130, 140, 150, or more features). In some embodiments, building the RT alignment model comprises scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file. In one embodiment, scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file comprises regression analysis (e.g., logistic regression). In one embodiment, scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file comprises linear determinant analysis (LDA). In some embodiments, scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file comprises, for each potential PSM, generating a score for each of the peptide target spectrum and the corresponding decoy spectrum and determining, based upon a difference between those scores, whether the peptide target spectrum has a PSM with the MS data file. In some embodiments, building the RT alignment model is an iterative process. In one embodiment, building the RT alignment model comprises using a first plurality of target spectra from the target library in combination with the MS data file to build a first RT alignment model, using a second plurality of target spectra from the target library in combination with the MS data file to build a second RT alignment model, and evaluating whether the first RT alignment model and the second RT alignment model are sufficiently convergent. In some embodiments, calculating the optimized RT window comprises using splines generated from a subset of PSMs in the first set of PSMs. In one embodiment, each PSM in the subset of PSMs meets a pre-selected false discovery rate (FDR) (e.g., 10%, 8%, 6%, 5%, 4%, 3% or less). In some embodiments, the subset of PSMs comprises at least 200 PSMs (e.g., at least 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 2000, 2500, 3000, or more PSMs). In some embodiments, the mass calibration curve is computed using linear regression. In one embodiment, the mass calibration curve is computed using a least squares regression to fit a polynomial. In one embodiment, the polynomial is a second order polynomial. In some embodiments, the searches in the series of searches of the recalibrated MS data file use an ascending set of values for mass tolerance. In some embodiments, a first search in the series of searches of the recalibrated MS data file uses a mass tolerance of 3 ppm. In one embodiment, subsequent searches in the series of searches use mass tolerances of 4 ppm, 5 ppm, 6 ppm, etc., up to 50 ppm. In some embodiments, the series of searches of the recalibrated MS data file is stopped when PSMs resulting from the series of searches reaches a maximum (e.g., a local maximum; or when the number of target counts does not increase after one, two, three, or more further searches). In some embodiments, scoring the target library against the recalibrated MS data file comprises performing regression analysis. In someWSGR Docket No. 53344-800.601embodiments, scoring the target library against the recalibrated MS data file comprises performing LDA. In some embodiments, scoring the target library against the recalibrated MS data file comprises generating feature vectors representative of each peptide target spectrum from the peptide target library, and scoring the feature vectors against the recalibrated MS data file. In one embodiment, each feature vector comprises at least 100 features (e.g., at least 120, 130, 140, 150 features; or about 160 to about 200 features, about 170 to about 190 features, or about 180 features). In some embodiments, scoring the target library against the recalibrated MS data file further comprises calculating precursor quantities for the set of optimized PSMs. In one embodiment, calculating a precursor quantity comprises summing detected fragment intensities. In one embodiment, summing the detected fragment intensities comprises summing at least the detected fragment intensities over the highest single scan. In some embodiments, summing the detected fragment intensities comprises summing the detected fragment intensities over all scans. In some embodiments, the machine learning algorithm is a neural network. In one embodiment, the neural network comprises a classifier having three fully-connected layers and a configurable hidden layer width using ReLU activation. In some embodiments, the neural network is trained using binary cross-entropy loss and multi-fold cross-validation. Alternatively or additionally, the method may further comprise obtaining a plurality of MS data files, each comprising retention times and corresponding m / z values, and searching each of the MS data files in the plurality of MS data files. In some embodiments, the plurality of MS data files comprises at least 100 MS data files (e.g., at least 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or more data files). In some embodiments, each MS data file in the plurality of MS data files is stored and / or searched in a single node (e.g., a single computer or a single processing unit in a distributed computing environment). In some embodiments, each MS data file in the plurality of MS data files is searched independently and / or wherein the MS data files in the plurality of MS data files are searched in parallel.

[0008] According to another aspect of the invention, a method for analyzing a mass spectrometry data file is provided, comprising processing a mass spectrometry data file using a library of predicted peptide spectra to identify i) a plurality of protein structure matches, ii) a first set of scores associated with the plurality of protein structure matches, and iii) a second set of scores associated with a plurality of decoy matches, dividing the first set of scores and the second set of scores into a plurality of blocks, each block containing a subset of the protein structure matches, computing, for each block, a count of the number of protein structure matches and a count of the number of decoy matches within that block, using the counts of the number of protein structure matches and the number of decoy matches within each block to estimate a false discovery rate (FDR) for each block, and aggregating the FDR estimates across the plurality of blocks to yield an overall FDR estimate for the plurality of protein structure matches. In some embodiments, the mass spectrometry data file is generated from a biological sample. In some embodiments, the method comprises calculating, for each block, an upper-bound q-valueWSGR Docket No. 53344-800.601estimate and a lower-bound q-value estimate based on the first set of scores and the second set of scores within that block. In some embodiments, the computing is performed using a distributed computing environment comprising a plurality of nodes. In some embodiments, the number of blocks is between 20,000 and 100,000. In one embodiment, the number of blocks is between 30,000 and 80,000. Alternatively or additionally, the method may further comprise selecting an acceptable relative error for quantization. In one embodiment, the acceptable relative error for quantization is less than 0.0005, 0.0004, 0.0003, 0.0002, or 0.0001.

[0009] According to another aspect of the invention, a non-transitory computer-readable storage media is provided, encoded with a computer program including instructions executable by one or more processors to perform operations comprising a method as described above.

[0010] According to another aspect of the invention, a system for performing quantitative DIA proteomics is provided, comprising a computer configured to perform the methods as described above, a data storage system for managing and processing one or more mzML files, each independently comprising data-independent analysis mass spectra of one or more biological samples, and a distributed computing environment (e.g., Apache Spark) for scalable execution of the data analysis pipeline. In some embodiments, the distributed computing environment is scalable to analyze at least 100 samples per day (e.g., at least 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, or more samples per day).INCORPORATION BY REFERENCE

[0011] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0013] FIG. 1 illustrates a block diagram providing an overview of an example embodiment of the methods for searching described herein. The block diagram illustrates steps in one implementation of the methods, variously described as a mass spectrometry search engine processing workflow. EachWSGR Docket No. 53344-800.601MS injection is processed separately, generating a set of putative PSMs with Linear Discriminant Analysis (LDA) and classifier scores.

[0014] FIG. 2 illustrates a block diagram of an example processing workflow described herein, termed the “MPF” workflow. In the search step, an example embodiment of a proteomics mass spectrometry search engine processing workflow provided herein is run on all files individually. Flexible downstream processing steps can be performed using a variety of methods or implementations as described herein.

[0015] FIG. 3 illustrates a workflow diagram of an example embodiment of methods provided herein including various modular implementations available in an MPF pipeline, provided herein. Such workflows may include open-source (but less-scalable) tools such as Mokapot and crema, as well as highly scalable Spark-based implementations (airpot, Cortado, Proffer). These are interchangeable, using well-defined interfaces and Apache Spark to communicate results between steps. A plug-in architecture allows for easy incorporation of new approaches or implementations; this permits rapid development and validation that is not possible with other tools. In certain examples provided herein, Cortado was employed to compute global PSM- and precursor-level q-values using an efficient bounded approximation of the MixMax algorithm, implemented with Apache Spark. This allows MPF to control global PSM FDR accurately and efficiently, with seamless scalability to thousands or tens of thousands of MS injections.

[0016] FIG. 4 illustrates evaluation of precursor count for an example embodiment of a search engine tested on diverse datasets in a single-node environment.

[0017] FIG. 5 illustrates search time for each respective dataset examined using the example search engine embodiment used in FIG. 4.

[0018] FIGs. 6A and 6B illustrate evaluation of cost per file and search time with respect to cohort size for an embodiment of a search engine provided herein.

[0019] FIG. 7 shows execution time for a mass spectrometry search engine search and MPF pipeline compared with a DIA-NN pipeline.

[0020] FIG. 8 illustrates evaluation of runtime and cost for an example embodiment of a search engine tested on diverse datasets in a cloud computing environment.

[0021] FIG. 9 illustrates evaluation of PSM, precursor, and protein group identifications from each of a plurality of searches of an example proteomics dataset using an example embodiment of a PSE provided herein.

[0022] FIG. 10 illustrates precision analysis of DIA-NN vs search using an example PSE provided herein on raw interpreted intensity for all precursors and shared precursors. DIA-NN was operated in “High Precision, Robust LC” mode on the “Hour Proteome” dataset.

[0023] FIG. 11 illustrates Venn diagrams of precursors measured in all replicates using an example PSE embodiment provided herein and DIA-NN.WSGR Docket No. 53344-800.601

[0024] FIG. 12 illustrates alignment of intensities for PSMs shared between search using an example PSE provided herein and DIA-NN. 10% of PSMs represented, selected at random, including alignment of observed retention times and residual difference.

[0025] FIG. 13 illustrates protein group analysis results from an Astral cohort study including a histogram of peptides per protein group using an example embodiment of a PSE provided herein.

[0026] FIG. 14 illustrates protein group completeness in a 37-sample study using an example embodiment of a PSE provided herein.

[0027] FIG. 15 illustrates precursor entrapment with a Shuffle Sequence FASTA concatenation using an alternate example embodiment of a PSE provided herein.

[0028] FIG. 16A illustrates an accuracy assessment of an example of a PSE provided herein with Matrix Matched Calibration Curves using example ground truth dataset to assess and compare search algorithm figures of merit. Analysis was done on a dataset of HeLa spiked into SILAC labeled HeLa. The top row illustrates a figures of merit quantile count relative to total entries at 1% FDR at specified LOD / Q thresholds. The middle row illustrates a precursor quantile at specified LOD / Q thresholds. The bottom row illustrates count distributions at specified LOD / Q thresholds for shared precursors between methods.

[0029] FIG. 16B illustrates an accuracy assessment of an alternate example of a PSE provided herein with Matrix Matched Calibration Curves using example ground truth dataset (QE MMCC) to assess and compare search algorithm figures of merit.

[0030] FIG. 17 illustrates an example extended assessment of an eye lens dataset according to some embodiments of methods described herein, including evaluation and identification of modified and unmodified peptides using aggregate identifications for each lens region.

[0031] FIG. 18 illustrates a comparison of precursor-level CV between methods.

[0032] FIG. 19A illustrates alignment of observed intensities between an example embodiment of a PSE provided herein and DIA-NN.

[0033] FIG. 19B illustrates agreement of detected retention types between a PSE of the invention and DIA-NN. While RTs for precursors without deamidations are consistent within one minute, deamidated precursors are detected with significantly greater RT residuals, especially in MBR mode. DIA-NN tends to detect deamidations at later RTs compared to the PSE, suggesting that it may be misidentifying unmodified peptides as deamidated.

[0034] FIG. 19C illustrates comparison of library dot products for all identifications. Skyline independently observed dot products for library free and MBR searches. Retention times observed in respective report files were used for alignment.

[0035] FIG. 19D illustrates alignment of intensities between data subsets between DIA-NN and an example embodiment of a PSE provided herein.WSGR Docket No. 53344-800.601

[0036] FIG. 20 illustrates an ECDF plot of library dot products for all identifications in an example dataset according to some embodiments described herein.

[0037] FIG. 21 illustrates a similar ECDF plot to that shown in FIG. 20 but focused on only deamidated PSMs.

[0038] FIG. 22 illustrates Distribution of spectral match quality assessed by Skyline (left) and measured intensity (right) for deamidated PSMs in the eye lens dataset.

[0039] FIG. 23 illustrates evaluation of identification counts at PSM-, precursor-, protein group, and one-peptide protein group (PG)-level, at 1% FDR, for a cancer data set.

[0040] FIG. 24 illustrates a Venn diagram of consistently identified precursor overlap in a comparison of an example embodiment of a PSE provided herein and DIA-NN across all replicates in an cancer study dataset.

[0041] FIGs. 25A, 25B, and 25C show an evaluation of a cancer case study. FIG. 25A shows nanoparticle (NP) precursor-resolved completeness for each workflow. FIG. 25B shows engine-reported normalized intensity by NP (using different nanoparticles; nanoparticle A “NPA” or nanoparticle B “NPB”), study group and workflow. FIG. 25C shows differential expression overlap comparison from NPA.

[0042] FIG. 26 illustrates significant differentially -abundant precursors identified for both nanoparticles in the cancer study of FIG. 25.

[0043] FIG. 27 illustrates the number of identified precursors for each of 2,561 RC injections, acquired over an eight-month period on four Astral instruments. Two different mixtures (“RC7” and “RC8”) were employed.

[0044] FIG. 28 shows a computer system that is programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION

[0045] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0046] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0047] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” orWSGR Docket No. 53344-800.601“less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0048] Certain inventive embodiments herein contemplate numerical ranges. When ranges are present, the ranges include the range endpoints. Additionally, every sub range and value within the range is present as if explicitly written out. The term “about” or “approximately” may mean within an acceptable error range for the particular value, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” may mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value may be assumed.

[0049] As used herein, “mass spectrometry search engine” (“MSE”) generally refers to an example implementation of methods described herein providing scalable search functionality for data-independent acquisition mass spectrometry data. Certain embodiments of methods provided herein may also be referred to as proteomics search engines (“PSEs”). Said terms are used interchangeably herein.

[0050] As used herein, modular pipeline framework (“MPF” or “PF”) generally refers to an example implementation of methods described herein providing a high-throughput proteomics workflow for protein discovery with reduced global false discovery rates. Certain embodiments of methods provided herein can also be referred to as scoring engine workflow pipelines (“SEWs”). Said terms are used interchangeably herein.

[0051] As used herein, the term “protein structure match” (also referred to as “peptide-spectrum match” or “PSM”) generally refers to a set comprising one or more associations between a mass spectrum and a peptide or protein sequence, indicating that the mass spectrum corresponds to the fragmentation pattern of that peptide or protein. A PSM may comprise one or more such associations. The term “false discovery rate” (FDR) generally refers to the expected proportion of false positives in a PSM. Lower FDR values indicate higher confidence in the validity of the PSMs, while higher FDR values suggest a greater likelihood of incorrect identifications.

[0052] The present invention provides a novel approach to peptide and protein identification and quantification in mass spectrometry, addressing the key challenges faced in large-scale proteomics studies. To address challenges in reducing false discovery rates (FDRs), the present application introduces data-independent analysis proteomics search engines which are efficient and highly scalable, designed to ensure robust identification and quantification of proteins while maintaining stringent FDR control across large-scale datasets. For example, by integrating linear discriminant analysis within the data-independent analysis search engines disclosed herein coupled with high-WSGR Docket No. 53344-800.601throughput analysis pipelines according to this disclosure, the accuracy, sensitivity, and scalability of proteomic analyses are enhanced, paving the way for advancements in research and clinical applications. This invention incorporates advanced algorithms optimized for sensitivity and reproducibility, allowing the example data-independent analysis search engine to achieve consistent performance even with extensive and diverse input data.

[0053] The invention provides proteomics search engines as well as modular processing pipelines. In some embodiments, provided is a DIA pipeline designed to support distributed search with cluster-scale post-processing. The search engine may be used for identification and scoring of peptide-spectrum matches (PSMs) in individual files, enabling efficient rapid parallel processing in a cloud-computing environment. Search results from the PSE are processed by a modular pipeline framework, which can be readily implemented in a distributed cluster computing environment (implemented, e.g. using Apache Spark). The MPF implements scalable versions of key workflow components - re-scoring, FDR estimation, protein inference and quantification, - and may support single-pass (“library-free”) and two-pass (“match between runs” or “MBR”) search workflows. The MPF may comprise a modular plug-in architecture to enable implementation of new workflows using existing modules or to add new modules. Each module can be distributed across a compute cluster, enabling the processing of extremely large experiments without high memory or storage requirement on any single node. For match between runs (MBR) processing, first-pass results may be filtered to 1% precursor FDR and a new library created, which is used for second pass searching.

[0054] The example data-independent analysis search engines according to this disclosure are designed to be integrated with example high-throughput analysis pipelines according to this disclosure, which complement the scoring from the example data-independent analysis search engines. Example workflows according to this disclosure are independently illustrated, for example, in FIGs. 1, 2, and 3. In some instances, steps in the exemplary workflows need not be performed exactly in the order presented and, optionally, may be performed simultaneously or in a partially overlapping manner. For example, in FIG. 1, the “Read MsFile” step may be performed at the same time as or prior to the “Read Library” step. Such workflows may be integrated with high-throughput proteomics workflows.

[0055] Methods for protein identification and quantification in mass spectrometry described herein may comprise the following steps: processing input data representing complex biological samples with numerous injections utilizing LDA to build a model that correlates predicted retention times and / or mass-to-charge ratios with empirical values, establishing a retention time window filter based on the LDA model, recalibrating mass-to-charge (m / z) values for all ions in a data file using average mass calibration curves derived from the processed batches.

[0056] Methods may further comprise optionally reporting or discarding PSMs that fall below a predetermined FDR threshold to ensure statistical validity to reduce an overall global FDR of theWSGR Docket No. 53344-800.601search engine. In this process, LDA may be used to processes small library batches to improve both accuracy and sensitivity in protein identification.

[0057] Methods of this disclosure may provide calibration and optimization techniques for calibrating mass tolerances and / or for extracting transition traces in MS data. This can further comprise: Processing small library batches at mass tolerance values ranging from 3 to 50 parts per million (ppm); utilizing LDA to train a neural network classifier based on the resulting PSMs, which enhances overall search accuracy and sensitivity; and / or providing global FDR estimates for PSMs, ensuring robust statistical validity throughout the analysis.

[0058] In some embodiments, the invention provides a method for processing mass spectrometry data, comprising receiving a plurality of retention times from a mass spectrometer; processing the plurality of retention times using a model configured to calibrate a retention time of the plurality of retention times; determining a first protein structure match based on the processed plurality of retention times; and training a machine learning model to output a second (or optimized) protein structure match, wherein the machine learning model is trained on adataset comprising the first protein structure match and the plurality of processed retention times.

[0059] A protein-structure match may comprise an identified peptide. In a proteomics context, peptides are obtained by proteolytic digestion of proteins in the input sample. Peptides are separated by liquid chromatography and ionized prior to mass spectrometry analysis. In a tandem mass spectrometry (MS / MS) experiment, peptides are selected for fragmentation based on their mass-to-charge (m / z) ratio, and the resulting fragment ions are detected to generate a mass spectrum. The peptide prior to the fragmentation step is referred to as a “precursor” or “precursor ion”.

[0060] Precursor peptides may be predicted based on the knowledge of the protein sequences in the sample and the protease used for digestion. For example, if trypsin is used, peptides may be predicted to be the sequences between lysine (K) and arginine (R) residues in the protein sequence, as trypsin cleaves at the C-terminal side of these residues. In some embodiments, Lys-C (lysyl endopeptidase) can be used for proteolytic digestion. In some applications, Lys-C cleaves peptide bonds after lysine (K) residues. In some embodiments, Lys-C exhibits greater specificity than trypsin. Where desired, Lys-C may be used in combination with trypsin. In some embodiments, chymotrypsin can be employed as a proteolytic enzyme. In some applications, chymotrypsin cleaves peptide bonds after large hydrophobic residues, which may include phenylalanine (F), tryptophan (W), tyrosine (Y), and leucine (L). In some embodiments, chymotrypsin provides coverage complementary to that obtained with trypsin. In some embodiments, Glu-C (V8 protease) can be used for peptide cleavage. In some applications, Glu-C cleaves peptide bonds after glutamic acid (E) and aspartic acid (D) residues. Where desired, Glu-C may be employed for alternative digestion strategies.WSGR Docket No. 53344-800.601

[0061] In some embodiments, pepsin can be used as a proteolytic enzyme. In some applications, pepsin exhibits non-specific cleavage patterns, where cleavage may occur with a tendency toward hydrophobic residues. In some embodiments, pepsin is active at low pH conditions. In some embodiments, Asp-N can be employed for proteolytic digestion. In some applications, Asp-N cleaves peptide bonds before aspartic acid (D) residues. In some embodiments, Asp-N generates peptides comprising N-terminal acidic residues.

[0062] The predicted peptides can be determined in silico from genomic or transcriptomic data. For example, predicted peptides can be generated by translating open reading frames (ORFs) from genomic or transcriptomic sequences. In some embodiments, predicted peptides can be obtained from public databases such as UniProt, NCBI RefSeq, or Ensembl. In some embodiments, predicted peptides can be derived from custom databases generated from specific organisms or sample types.

[0063] Input data suitable for processing by the methods described herein include, but are not limited to, mass spectrometry files comprising retention times, mass-to-charge (m / z) values, intensities, ion mobility values, and / or other relevant spectral data. Such files may be in formats including, but not limited to, mzML, mzXML, RAW, Parquet, or other proprietary or open formats used in mass spectrometry data storage. A data file may correspond to a single mass spectrometry acquisition or injection, or may comprise multiple acquisitions or injections. In some embodiments, the data file corresponds to a single injection. In some embodiments, the input data is loaded from a storage medium (including, but not limited to, local disk storage, network-attached storage, cloud storage services) into a computing environment for processing. The computing environment may comprise one or more processing units (e.g., CPUs, GPUs) and memory resources suitable for handling datasets typical of mass spectrometry experiments. Data files may be processed in parallel. Files may be loaded in a single thread or in multiple threads. In some embodiments, “lazy loading” is used, where data is accessed only as needed. Each file may be processed independently, allowing for scalable execution in a distributed computing environment, such as a cloud computing platform. This parallel processing capability enables efficient handling of large datasets, reducing overall processing time and resource requirements.

[0064] The data may be converted into an internal representation suitable for processing by the methods described herein. This may involve parsing the file format, extracting relevant spectral data, and organizing it into data structures optimized for efficient access and computation. In some embodiments, the data is organized in a database which may be processed in memory or stored.

[0065] In some embodiments, the database comprises a tree structure. For example, the database may be organized as an R-tree. An R-tree refers to a spatial indexing data structure designed for efficiently organizing and querying multi-dimensional data. It is a tree structure that groups nearby objects using their bounding rectangles and organizes them hierarchically. In an R-tree, each node contains a bounding box that encloses all the bounding boxes of its child nodes. This structureWSGR Docket No. 53344-800.601allows for efficient queries given parameters along multiple dimensions, such as range searches and nearest neighbor searches, by quickly eliminating large portions of the dataset that do not intersect with the query region. Dimensions relevant to mass spectrometry data that may be indexed in an R-tree include m / z (mass-to-charge ratio), retention time, intensity, and ion mobility. For example, an R-tree can help to efficiently locate all spectra or peaks within a specified m / z and retention time range, which is particularly useful in data-independent acquisition (DIA) workflows where large volumes of spectral data need to be queried rapidly. R-trees may also be used for chromatographic peak detection, data compression, and real-time spectral library matching. Performance benefits include faster query times for range searches and nearest neighbor searches, reduced memory usage by loading only relevant portions of the data, and scalability to handle large datasets typical in mass spectrometry experiments.

[0066] Any suitable R-tree structure may be used, including original R-trees and packed R-trees. Specifically, original R-trees may be well-suited for dynamic datasets with frequent insertions and deletions, while packed R-trees may be preferred for static datasets where bulk loading can enhance query performance. In original R-tree structures, nodes are inserted one at a time as the data arrives, and the tree is grown incrementally. In packed R-tree structures, the data is loaded using bulk-loaded algorithms, and the tree is built bottom-up from pre-sorted data to minimize overlap and dead space between bounding rectangles.

[0067] In some embodiments, decoy sequences are generated for each target in the library. In some embodiments, a decoy sequence is a reversed, shuffled, or mutated version of the target sequence having a different predicted mass spectrum than the corresponding target sequence. In some embodiments, decoys are generated by mutating the target sequence at specific positions. In the case of peptide targets, one exemplary decoy generation approach is to introduce amino acid substitutions, for example, at the second and penultimate amino acid residue positions. In some applications, a mutation pattern is applied according to a substitution key, for example ACDEFGHIKLMNPQRSTUVWY LSEDLLSVLVLQLNLTSSLLS. In other words, for a given position in a peptide target, an alanine (“A”) residue at that position can be replaced with a leucine (“L”) residue at the same position in the decoy sequence; a cysteine (“C”) residue at that position can be replaced with a serine (“S”) residue at the same position in the decoy sequence; an aspartic acid (“D”) residue at that position can be replaced with a glutamic acid (“E”) residue at the same position in the decoy sequence; etc. Predicted fragment ion masses in the decoy mass spectrum can be adjusted according to the corresponding mass shifts. In some embodiments, decoys are generated using sequence and spectrum manipulation techniques such as those employed in DIA -NN.

[0068] Peptide target spectra and the corresponding decoy spectra may be represented by multidimensional feature vectors. Each feature vector may comprise dimensions such as m / z values, retention times, intensities, ion mobility values, and other relevant spectral characteristics. TheWSGR Docket No. 53344-800.601feature vectors can be stored in the R-tree structure for efficient querying and retrieval during the search process. In some embodiments, the multi-dimensional feature vectors may comprise at least 50 features (e.g., at least 75, 100, 110, 120, 130, 140, 150, or more features). In other embodiments, the multi-dimensional feature vectors may comprise at least 100 features (e.g., at least 120, 130, 140, 150, or more features; or about 160 to about 200 features, about 170 to about 190 features, or about 180 features). The number of features relates to the dimensionality of the vectors used to represent the spectra, and features may be dependent or independent. Exemplary features include, but are not limited to predicted retention time, the width of the retention time window, predicted number of fragments, mass of each predicted fragment, mass of the 2, 5, 7, or 10 largest / most intense fragments (summed), mass of all fragments (summed), peak intensity for each fragment, average peak intensity for the fragments, average peak intensity for the X largest / most intense fragments, and amino acid composition features such as the number of Ala or Gly residue (as well as similar features for all 20 amino acids, as well as post-translationally modified amino acids).

[0069] Starting from the input data, RT prediction and / or mass calibration curves may be generated. In some applications, subsets of the database may be processed in batches. In some embodiments, a subset of peptide target spectra corresponds to about 0.1% to about 25% of the targets in the target library (e.g., up to about 20%, 15%, 10%, 7%, 5%, 4%, 3%, 2%, or 1% of peptide targets in a peptide target library). In some embodiments, the subset of peptide target spectra comprises about 20,000 peptide target spectra (e.g., at least about 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000 or more peptide target spectra).

[0070] In some embodiments, a retention time (RT) window filter to be used is determined, for example from an RT alignment model. The model configured to calibrate the retention time may comprise a retention time (RT) alignment model and the RT window filter. In some applications, a regression model can be used, for example an LDA model as described herein. In some embodiments, the model configured to calibrate the retention time uses a dataset comprising empirical retention data and predicted retention data. In some embodiments, the predicted retention data is obtained from a target library. For example, the target library may comprise predicted retention times for peptides based on their sequences, which can be generated using machine learning models trained on empirical retention time data. In some embodiments, such empirical retention time data is obtained from prior mass spectrometry experiments. For example, empirical retention times can be derived from previously acquired mass spectrometry datasets where peptides have been confidently identified and their retention times recorded. Besides a retention time, the target library may comprise a plurality of mass-to-charge (m / z) values, wherein each retention time of the plurality of retention times is paired with a corresponding m / z value of the plurality of m / z values; intensities for each m / z value; corresponding ion mobility data (in instances where ionWSGR Docket No. 53344-800.601mobility is used to separate ions prior to their detection in a mass spectrometer); and / or other relevant data.

[0071] As part of constructing an RT alignment model, a scoring function may be used to score peptide-spectrum matches (PSMs) based on features derived from the mass spectrometry data. The scoring function may incorporate features such as mass accuracy, fragment ion intensities, retention time deviations, and other relevant spectral characteristics. The scoring function may comprise regression analysis (e.g., logistic regression) or linear discriminant analysis (LDA). For example, the scoring produces a score for each vector (e.g. a projection along a single axis) that allows targetdecoy vector pairs to be quickly compared to one another. Target peptide spectra that score well compared to their corresponding decoy spectra can be identified as a PSM.

[0072] In some implementations, the RT prediction and calibration process may be repeated in an iterative fashion to refine the RT alignment model and / or RT window filter. For instance, a first plurality of target spectra from the target library in combination with the MS data file is used to build a first RT alignment model, followed by using a second plurality of target spectra from the target library in combination with the MS data file to build a second RT alignment model, and evaluating whether the first RT alignment model and the second RT alignment model are sufficiently convergent.

[0073] The RT window filter may depend on the chromatographic conditions in the experimental acquisition, the quality of the RT information in the provided library (e.g. empirical vs. high-quality predictions and low-quality predictions), etc. The RT window filter may be determined using a specified number of PSMs at a given FDR (for example, 1000 PSMs at 5% FDR). For example, using a set of PSMs at a given FDR, (e.g., 1000 PSMs at 5% FDR), splines (e.g., piecewise-defined functions, often piecewise-defined polynomials) can be fit to the RT data of the set of PSMs. The splines resulting from the analysis can be (or can be used to generate) the RT window filter.

[0074] Mass spectrometry data may be recalibrated using the computed mass calibration curve. The recalibration may be performed using a least squares method to fit a polynomial of configurable order, such as a 2nd order polynomial. The ppm tolerance (referring to the difference between observed and expected retention times expressed in ppm) of the RT alignment model can be optimized, for example using LDA. The optimization may include processing a small batch of library targets at a certain value range, such as ranging from 3 to 50 ppm. The processing can be performed in ascending fashion (e.g. subsequent searches in the series of searches use increasingly larger mass tolerances, such as 4 ppm, 5ppm, 6 ppm, etc., up to 50 ppm), and may halt when target counts do not increase above a maximum after a number of further iterations, such as one, two, three, four, or more iterations. In some embodiments, a subset (or small batch) of library targets corresponds to about 0.1% to about 25% of the peptide targets in a peptide target library (e.g., up to about 20%, 15%, 10%, 7%, 5%, 4%, 3%, 2%, or 1% of the peptide targets in the peptide targetWSGR Docket No. 53344-800.601library). In some embodiments, a subset (or small batch) of library targets comprises about 20,000 peptide target spectra (e.g., at least about 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000 or more peptide target spectra).

[0075] Following optimization of the ppm tolerances, the full dataset can be scored. Such scoring may be performed using LDA. Similar to the batch scoring used to generate the RT alignment model, scoring the full dataset (e.g., the target library) can employ feature vectors, with the feature vectors representing corresponding peptide targets (or peptide target spectra) in the target library. The feature vectors can be multi-dimensional. In certain embodiments, the multi-dimensional feature vectors may comprise at least 120 features (e.g., at least 130, 140, 150, or more features; or about 160 to about 200 features, about 170 to about 190 features, or about 180 features). In some applications, precursor quantities can be computed. The computation may include summing detected fragment intensities over one or more scans. For example, the one or more scans can include a peak scan (i.e., a scan having the highest sum of all utilized fragment intensities). As another example, the one or more scans can include the peak scan plus all other scans having a summed fragment intensity of all utilized fragments that is greater than a threshold. The threshold can be, for example, 50%, 55%, 60%, 65%, 70%, 75%, 80%, or 85%, in some embodiments, 65%, of the sum of all utilized fragment intensities in the peak scan.

[0076] Based on the processed plurality of retention times, a protein structure match (e.g. an optimized PSM) may be determined. The determination may be based on the processed plurality of retention times (e.g., the RT alignment model and / or RT window filter) and the recalibrated plurality of corresponding m / z values. The determination may further use a plurality of signal intensities, wherein each retention time of the plurality of retention times is paired with a corresponding signal intensity of the plurality of signal intensities. The determination may also use a plurality of ion mobility values, wherein each retention time of the plurality of retention times is paired with a corresponding ion mobility value. An ion mobility calibration curve may be computed, which can then be used recalibrating the received plurality of corresponding ion mobility values using the computed ion mobility calibration curve. The determination of the protein structure match may further be based on the recalibrated plurality of corresponding ion mobility values. Signal intensities and ion counts may be normalized as necessary prior to use in the determination of the first protein structure match.

[0077] A PSM (e.g., an optimized PSM) may comprise a feature vector and / or a feature vector score (e.g., a single-value score) and, optionally, related meta-data. The feature vector may comprise a plurality of features derived from the mass spectrometry data, such as m / z values, retention times, intensities, ion mobility values, and other relevant spectral characteristics as described herein. The associated meta-data may include a sequence for the matched peptide, a retention time for the matched peptide, ion mobility time for the matched peptide, and other information.WSGR Docket No. 53344-800.601

[0078] The dataset comprising the protein structure matches and the plurality of processed retention times may be used to train a classification system to output a second (or optimized) protein structure match. The classification system may be trained to classify candidate protein structure matches as either true or false matches based on features derived from the mass spectrometry data, including but not limited to retention time, mass-to-charge ratio, intensity, fragmentation patterns, and other relevant spectral characteristics. The classification system may output a score or probability indicating the likelihood that a given candidate match is a true positive. This score can then be used to rank and filter the candidate matches, allowing for the selection of high-confidence protein structure matches while controlling the false discovery rate (FDR).

[0079] In various embodiments, the classification system may employ one or more machine learning models adapted for classification tasks. Suitable models include, but are not limited to, support vector machines (SVM) with linear or non-linear kernels, decision tree classifiers, random forest ensembles, gradient boosting machines, naive Bayes classifiers, k-nearest neighbors (k-NN) algorithms, and logistic regression models. In particular embodiments, neural network -based classifiers may be utilized, including feedforward neural networks, convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTM) networks, gated recurrent units (GRUs), transformer-based architectures, and deep residual networks (ResNets). The neural network classifiers may comprise one or more hidden layers, wherein each layer includes a plurality of nodes with associated weights and biases that are optimized during a training phase using backpropagation and gradient descent methods.

[0080] In some embodiments, the classifier used is a neural network classifier. In some embodiments, a neural network classifier is trained using PSMs passing a threshold FDR or a threshold number of PSMs. In some applications, the threshold FDR can be approximately 30%, 40%, 50%, or 60% FDR. In some applications, the threshold number of PSMs can be approximately 50,000 PSMs. In one embodiment, the neural network classifier comprises three fully -connected layers with ReLU activation. In some embodiments, the hidden layer width is configurable. In some embodiments, the neural network classifier is trained using binary cross-entropy loss. In some embodiments, a multi-fold cross-validation scheme is used during training. In some applications, the multi-fold cross-validation scheme ensures that each PSM is not used to train the network that scores it. In some embodiments, cross-fold normalization is applied to ensure well-calibrated scores.

[0081] In some embodiments, inference is performed with a neural network classifier to output a second (or optimized) protein structure match. In some embodiments, the neural network classifier outputs a score for each candidate protein structure match, indicating the likelihood that the match is correct. The scores may be used to rank the candidate matches, allowing for the selection of high-confidence matches while controlling the false discovery rate (FDR). For example, after neural network training, PSM FDR estimates may be recomputed based on the neural network classifierWSGR Docket No. 53344-800.601scores. Such FDR estimates may be referred to as an “optimized FDR”. In some embodiments, results are written to a tabular output format, for example a Parquet file. In some applications, a user-configurable FDR threshold may be applied to filter results. Output results may be subsequently displayed to the user.

[0082] In some embodiments, the method comprises determining a plurality of first protein structure matches, wherein the machine learning model is trained on the plurality of first protein structure matches, the plurality of processed retention times, and, optionally, the recalibrated plurality of corresponding m / z values and / or the plurality of corresponding signal intensities and / or the recalibrated plurality of corresponding ion mobility values. Each protein structure match of the plurality of first protein structure matches may have an FDR that is below a pre-selected threshold. For example, the pre-selected FDR threshold for the first protein structure matches is 50% FDR or less, such as 45%, 40%, 35%, 30%, or 25%. In some applications, the method further comprises outputting a list comprising a plurality of second protein structure matches, wherein each second protein structure match passes a pre-selected FDR threshold. In some embodiments, the pre-selected FDR threshold for the plurality of second protein structure matches is user-selected. In some applications, the pre-selected FDR threshold for the plurality of second protein structure matches is an FDR of 1% or less.

[0083] The exemplary data-independent analysis search engines described herein may rely on the implementation of linear discriminant analysis (LDA). This statistical method is employed to build models that correlate predicted retention times with empirical values, thereby enhancing the accuracy of retention time window filtering and mass calibration. By utilizing LDA in this manner, the example data-independent analysis search engines significantly improve the precision of peptide-spectrum match (PSM) identification and quantification. Linear discriminant analysis (LDA) is a statistical technique used for dimensionality reduction and classification. In the context of mass spectrometry-based proteomics, LDA can be employed to enhance the identification and quantification of peptides and proteins by effectively separating true positive identifications from false positives based on multiple features derived from the mass spectrometry data. Generally, LDA operates by taking multiple input features, such as various spectral match quality metrics in proteomics applications, and analyzing how these features can be combined to distinguish between different classes of data. The technique works by identifying a linear combination of these input features that achieves a best possible separation between two or more classes, for example to distinguish between target and decoy matches in peptide identification. This can be accomplished by projecting the multi-dimensional data onto a lower-dimensional space, often reducing it to a much lower number of dimensions, where the separation between classes is maximized for optimal discriminatory power. This approach makes LDA particularly effective for classification tasks whereWSGR Docket No. 53344-800.601multiple correlated features are integrated into a single, interpretable discriminant score that can reliably differentiate between competing hypotheses or classifications.

[0084] In some embodiments, a method comprises receiving, at a plurality of processing units, mass spectrometry data (e.g. data produced by at least one mass spectrometer); processing the mass spectrometry data, wherein the processing comprises calibrating a retention time of the mass spectrometry data; and determining a plurality of protein structure matches.

[0085] Processing the mass spectrometry data may comprise performing the methods described herein, such as the Peptide Search Engine methods. The plurality of processing units may be part of a distributed computing environment, such as a cloud computing platform. The mass spectrometry data may be processed in parallel across the plurality of processing units, allowing for scalable execution and efficient handling of large datasets. Each processing unit may independently process a subset of the mass spectrometry data, enabling concurrent execution and reducing overall processing time. For example, the plurality of processing units comprises at least two nodes. As used herein, the term node refers to a single computing device or server within a distributed computing environment. Each node may comprise one or more processing units (e.g., CPUs, GPUs) and memory resources suitable for handling data typical of mass spectrometry experiments. Nodes may be interconnected via a network, allowing for communication and data exchange between them during parallel processing tasks.

[0086] The modular pipeline frameworks described herein are capable of directing processing of mass spectrometry data (as discussed above and elsewhere herein) and scaling to handle large datasets typical of mass spectrometry experiments. For example, the MPFs of the invention may process at least 1000 or more data files (e.g., at least 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10,000, 11,000, 12,000, 13,000, 14,000, 15,000, 20,000, 25,000, or more data files) comprising the plurality of mass spectra within a 24-hour period. The MPFs may process the mass spectrometry data while maintaining a global false discovery rate (FDR) below a predetermined threshold.

[0087] In some embodiments, the MPFs processes a population-scale proteomics study while maintaining a global false discovery rate (FDR) below a predetermined threshold. For example, the MPFs may maintain a global FDR of less than about 50% (e.g., less than about 45%, 40%, 35%, or 30%) while processing a population-scale proteomics study. In some embodiments, the MPFs maintains a global FDR of less than about 5% (e.g., less than about 4%, 3%, 2%, 1.5%, or 1%) while processing a population-scale proteomics study.

[0088] In some embodiments, the invention provides method for estimating the false discovery rate among peptide-spectrum matches. In some embodiments, a mixture-maximum (mix-max) FDR estimation procedure is used. In one aspect, the method comprises obtaining target PSM scores and corresponding decoy PSM scores from a database search of tandem mass spectra. The method mayWSGR Docket No. 53344-800.601further comprise partitioning false discoveries into (a) a first component comprising false positives arising from foreign spectra not generated by any peptide in the target database, and (b) a second component comprising false positives arising from native spectra wherein the true generating peptide exists in the database but an incorrect match is reported. The method may further comprise estimating the first component by computing the product of the estimated proportion of foreign spectra and the count of target PSMs exceeding a score threshold. The method may further comprise estimating the second component by computing, for each native spectrum, the probability that its best-scoring target match derives from random matching rather than correct identification, using the decoy score distribution as a proxy for the null distribution. The method may further comprise aggregating said component estimates to yield an unbiased FDR estimate without discarding high-scoring correct identifications. The method may provide accurate FDR estimates while preserving sensitivity, overcoming limitations of prior art target-decoy competition protocols that randomly eliminate valid identifications.

[0089] In some embodiments, the methods for estimating the false discovery rate comprise dividing score data into a plurality of blocks, each block containing a subset of PSMs. The number of PSMs may be predetermined. In some embodiments, the number of blocks is between 20,000 and 100,000, or between 30,000 and 80,000. The method may also comprise selecting an acceptable relative error for quantization, for example, <0.0001. Parameters may be selected to balance calculation speed and quantization granularity. Score space quantization may be performed, for example, using Spark’s QuantileDiscretizer. The method may further comprise computing, for each block, the counts of target and decoy PSMs within that block. The method may further comprise calculating upper-bound and lower-bound q-value estimates for each block based on the target and decoy counts, accounting for tied scores within blocks. The method may further comprise assigning q-values to individual PSMs based on their respective blocks. In some embodiments, results may bound the non-quantized MixMax result, and upper and lower bounds can be identical when one PSM is present per block. Performance characteristics in some applications include computational complexity that is determined by the number of blocks, where runtime can be constant for any PSM count given a fixed block number. Quantization and count aggregation complexity can grow sub-linearly with PSM count. Calculations can be easily distributed and can efficiently handle thousands of MS acquisitions in some implementations. In some embodiments, Posterior Error Probability (PEP) estimation can be performed. The proportion of decoys within each block can be computed, and a rolling average filter with configurable width may be applied. A cumulative maximum can be taken to provide a monotonic score-PEP relationship.

[0090] In some applications, result assignment may be performed where results are assigned to PSMs within each block without interpolation. Upper-bound q-values, lower-bound q-values, and PEPs can be reported. Downstream analyses can use the upper-bound q-value for conservativeWSGR Docket No. 53344-800.601filtering in some embodiments, where little practical difference may exist between upper and lower bound estimates.

[0091] In some embodiments, other FDR calculations can be performed. Global precursor FDR can follow the same process as PSM FDR, where the best-scoring PSM for each precursor may be selected. Protein-level FDR can also be computed in some applications.

[0092] Mass spectroscopy data files may be generated by analysis of a biological sample, including, e.g., a complex biological sample. The methods of the invention are suitable for analyzing data files generated from a diverse range of biological samples and run conditions, including a range of different sample types, sample preparation techniques, chromatographic conditions and mass spectroscopy instrument type.

[0093] The present disclosure provides a range of samples that can be assayed using the methods provided herein. A sample may be a biological sample (e.g., a sample derived from a living organism). A sample may comprise a cell or be cell-free. A sample may comprise a biofluid, such as blood, serum, plasma, urine, or cerebrospinal fluid (CSF). Samples of the present disclosure include biological samples from a subject. A method may include analyzing a sample from a single subject, or analyzing samples from multiple subjects. The subject may be a human or a non -human animal. The biological samples can contain a plurality of proteins or proteomic data, which may be analyzed after adsorption of proteins to the surface of the various sensor element (e.g., particle) types in a panel and subsequent digestion of protein coronas. Proteomic data can comprise nucleic acids, peptides, or proteins. A biofluid may be a fluidized solid, for example a tissue homogenate, or a fluid extracted from a biological sample. A biological sample may be, for example, a tissue sample or a fine needle aspiration (FNA) sample. A biological sample may be a cell culture sample. For example, a biofluid may be a fluidized cell culture extract.

[0094] A wide range of samples are compatible for use within the methods and compositions of the present disclosure. The biological sample may comprise plasma, serum, urine, cerebrospinal fluid, synovial fluid, tears, saliva, whole blood, a blood component (e.g., plasma or white blood cells), milk, nipple aspirate, ductal lavage, vaginal fluid, nasal fluid, ear fluid, gastric fluid, pancreatic fluid, trabecular fluid, lung lavage, sweat, crevicular fluid, semen, prostatic fluid, sputum, fecal matter, bronchial lavage, fluid from swabbings, bronchial aspirants, fluidized solids, fine needle aspiration samples, tissue homogenates, lymphatic fluid, cell culture samples, or any combination thereof. The biological sample may comprise blood or a blood component. The biological sample may comprise multiple biological samples (e.g., pooled plasma from multiple subjects, or multiple tissue samples from a single subject). The biological sample may comprise a single type of biofluid or biomaterial from a single source. A biological sample may comprise a nerve biopsy.

[0095] Various methods of the present disclosure utilize blood or blood components (e.g., red blood cells, buffy coats, plasma). Contrasting many tissue biopsies, which can be damaging and costWSGR Docket No. 53344-800.601intensive, blood collection is often relatively facile and benign, and is therefore suitable for routine and low-risk patient monitoring. Furthermore, as human blood is estimated to contain over 5000 types of protein groups whose abundances and forms (e.g., post-translationally modifications and variant types) can be responsive to , the blood proteome offers a biological state changes are often evidenced by subtle changes in blood protein composition. A method of the present disclosure may use whole blood (e.g., untreated blood drawn from a subject). A method of the present disclosure may also use a treated or partitioned blood sample. In some cases, a sample comprises plasma, buffy coat, white blood cells, platelets, hematocrit, red blood cells, serum, blood clots or any combination thereof. In some cases, plasma, buffy coat, white blood cells, platelets, hematocrit, red blood cells, serum, blood clots or any combination thereof are extracted from a blood sample for use in a method disclosed herein.

[0096] In some cases, a method utilizes serum. As used herein, “serum” may denote the liquid fraction remaining after a blood sample clots. As a blood sample left at room temperature will typically clot within 15-60 minutes, serum may be prepared by incubating a blood sample at or above room temperature, for example at 25 °C or at 37 °C, respectively. After at least about 10 minutes, at least about 15 minutes, at least about 20 minutes, at least about 30 minutes, at least about 40 minutes, at least about 50 minutes, or at least about 60 minutes, the blood clots may be separated from solution through centrifugation. While serum is often prepared non -hemo lyzed (e.g., wherein blood cells remain intact through clotting and removal), some methods of the present disclosure may utilize serum derived from hemolyzed blood samples.

[0097] In some cases, a method utilizes plasma. As used herein, “plasma” may denote a fraction collected from blood pretreated with an anticoagulant and separated from blood cells and platelets. Contrasting with serum, plasma typically contains an array of clotting factors, such as fibrinogen, prothrombin, and proaccelerin. As the concentrations and forms of these species can reflect certain health conditions, plasma analysis can provide greater diagnostic insight than serum analysis for some biological states. Plasma samples can be prepared treating blood with an anticoagulant, and then centrifuging the treated blood. The anticoagulant may comprise citrate, ethylenediaminetetraaceticacid (EDTA), potassium oxalate, hirudin, argatroban, ximelagatran, heparin, fondaparinux, or any combination thereof.

[0098] Centrifugation parameters affect the proteins which remain in solution, and therefore may be modified depending on the biomolecules of interest for detection from plasma or serum.Centrifugation may be performed for at least 2 minutes, at least 4 minutes, at least 6 minutes, at least 8 minutes, at least 10 minutes, at least 12 minutes, at least 15 minutes, at least 20 minutes, or at least 30 minutes. Centrifugation may be performed for at most 30 minutes, at most 20 minutes, at most 15 minutes, at most 10 minutes, at most 8 minutes, at most 6 minutes, at most 4 minutes, or at most 2 minutes. Centrifugation may impart at least 100 gravitational force equivalents (g), at least 200 g, atWSGR Docket No. 53344-800.601least 300 g, at least 400 g, at least 500 g, at least 600 g, at least 800 g, at least 1000 g, at least 1200 g, at least 1500 g, at least 1800 g, at least 2000 g, at least 2500 g, at least 3000 g, at least 4000 g, at least 5000 g, at least 6000 g, at least 8000 g, or at least 10000 g. The centrifugation may impart at most 100 g, at most 200 g, at most 300 g, at most 400 g, at most 500 g, at most 600 g, at most 800 g, at most 1000 g, at most 1200 g, at most 1500 g, at most 1800 g, at most 2000 g, at most 2500 g, at most 3000 g, at most 4000 g, at most 5000 g, at most 6000 g, at most 8000 g, or at most 10000 g.

[0099] The biological sample may be diluted or pre-treated. The biological sample may undergo depletion (e.g., albumin removal from serum or plasma) prior to or following contact with a particle or plurality of particles. The biological sample may also undergo physical (e.g., homogenization or sonication) or chemical treatment prior to or following contact with a particle or plurality of particles. The biological sample may be diluted prior to or following contact with a particle or plurality of particles. The dilution medium may comprise buffer or salts, or be purified water (e.g., distilled water). Different partitions of a biological sample may undergo different degrees of dilution. A biological sample or a portion thereof may undergo a 1.1 -fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 8-fold, 10-fold, 12-fold, 15-fold, 20-fold, 30-fold, 40-fold, 50-fold, 75-fold, 100-fold, 200-fold, 500-fold, or 1000-fold dilution. For example, a plasma sample may be subjected to a 5-fold dilution with buffer prior to analysis.

[0100] The compositions and methods of the present disclosure can be used to measure, detect, and identify specific proteins from biological samples. Examples of proteins that can be identified and measured include highly abundant proteins, proteins of medium abundance, and low-abundance proteins.

[0101] Prior to analysis by mass spectroscopy, biological samples may be subjected to various sample preparation techniques, such as protein extraction, reduction and alkylation of disulfide bonds, enzymatic digestion (e.g., with trypsin, Lys-C, chymotrypsin, Glu-C, pepsin, Asp-N), and peptide cleanup (e.g., using solid-phase extraction). The resulting peptide mixtures can then be analyzed by liquid chromatography-tandem mass spectrometry (LC-MS / MS) to generate the MS data files that are processed by the methods described herein. The methods are designed to handle the variability introduced by different sample preparation techniques and chromatographic conditions, making them broadly applicable to a wide range of proteomics studies.

[0102] Various chromatographic conditions may be used, such as different column types (e.g., Cl 8, C8, HILIC), column lengths (e.g., 15 cm, 25 cm, 50 cm), and gradient lengths (e.g., 30 min, 60 min, 120 min). The methods are robust to these variations, allowing for accurate protein identification and quantification across diverse experimental setups. Additionally, the methods can be applied to MS data files generated from various types of mass spectrometers, including but not limited to Orbitraps, time-of-flight (TOF) instruments, quadrupole mass analyzers, and ion traps. This versatility ensuresWSGR Docket No. 53344-800.601that the methods can be adopted for analyzing complex biological samples under a variety of conditions.

[0103] Aspects of the present disclosure provide compositions, systems, and methods for collecting biomolecules on particles (e.g., nanoparticles and microparticles) (as well as other types of sensor elements such as polymer matrices, filters, rods, and extended surfaces). In some cases, a particle may adsorb a plurality of biomolecules upon contact with a biological sample, thereby forming a biomolecule corona on the surfaces of the particles. In some cases, the biomolecule corona may comprise proteins, lipids, nucleic acids, metabolites, saccharides, small molecules (e.g., sterols), and other biological species present in a sample. In some cases, a biomolecule corona comprising proteins may also be referred to as a ‘protein corona’, and may refer to all constituents adsorbed to a particle (e.g., proteins, lipids, nucleic acids, and other biomolecules), or may refer only to proteins adsorbed to the particle.

[0104] A plurality of particles may be contacted with a biological sample comprising proteins, wherein each particle adsorbs a plurality of proteins from the biological sample to its surface. The different particles may represent distinct particle types, such that each particle differs from the other particles by at least one physicochemical property. This difference in physicochemical properties can lead to the formation of different protein corona compositions on the particle surfaces.

[0105] The composition of the biomolecule corona may depend on a property of the particle. In many cases, the composition of the biomolecule corona is strongly dependent on the surface of the particle. Characteristics such as particle surface material (e.g., ceramic, polymer, metal, metal oxide, graphite, silicon dioxide, etc.), surface texture (rough, smooth, grooved, etc.), surface functionalization (e.g., carboxylate functionalized, amine functionalized, small molecule (e.g., saccharide) functionalized, etc.), shape, curvature, and size can each independently serve as determinants for biomolecule corona composition. In addition to surface features, the particle core composition, particle density, and particle surface area to mass ratio may each influence biomolecule corona composition. For example, two particles comprising the same surfaces and different cores may form different biomolecule coronas upon contact with the same sample.

[0106] Biomolecule corona formation may also be influenced by sample composition. For example, a first sample condition (e.g., low salinity) might favor the solubility of a particular analyte (e.g., an isoform of Bone Morphogenic Protein 1 (BMP1)), and thereby disfavor its binding in a biomolecule corona, while a second sample condition (e.g., high salinity) may diminish the solubility of the analyte, thereby driving its incorporation into a biomolecule corona.

[0107] Biomolecule corona composition may also depend on molecular level interactions between the biomolecules, themselves. An energetically favorable interaction between two biomolecules may promote their co-incorporation into a biomolecule corona. For example, if a first protein adsorbed to a particle comprises an affinity for a second protein in solution, the first protein may bind to aWSGR Docket No. 53344-800.601portion of the second protein, thereby driving its binding to the particle or to other proteins of the biomolecule corona of the particle. Analogously, a first biomolecule disposed within a biomolecule corona may comprise an energetically unfavorable interaction with a second biomolecule in a biological sample, thereby disfavoring its incorporation into a biomolecule corona. In part owing to these inter-biomolecule dependencies, biomolecule coronas provide sensitive platforms for directly and indirectly sensing biomolecules from a biological sample.

[0108] Biomolecules collected on a particle may be subjected to further treatment prior to mass spectroscopy. A method may comprise collecting a biomolecule corona or a subset of biomolecules from a biomolecule corona. The collected biomolecule corona or the collected subset of biomolecules from the biomolecule corona may be subjected to further particle-based analysis (e.g., particle adsorption). The collected biomolecule corona or the collected subset of biomolecules from the biomolecule corona may be purified or fractionated (e.g., by a chromatographic method). The collected biomolecule corona or the collected subset of biomolecules from the biomolecule corona may be analyzed (e.g., by mass spectrometry). Furthermore, as biomolecule corona composition is dependent on solution-phase and particle-bound biomolecules as well as sample conditions (e.g., pH, osmolarity, lipid concentration), biomolecule corona composition can provide a sensitive measure of biomolecules which are not bound to a particle and of sample conditions.

[0109] The particles and methods of use thereof disclosed herein can bind a large number of unique biomolecules (e.g., proteins) in a biological sample (e.g., a biofluid). For example, a particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising at least 5 protein groups, at least 10 protein groups, at least 15 protein groups, at least 20 protein groups, at least 25 protein groups, at least 50 protein groups, at least 80 protein groups, at least 100 protein groups, least 150 protein groups, at least 180 protein groups, at least 200 protein groups, at least 250 protein groups, at least 300 protein groups, at least 350 protein groups, at least 400 protein groups, at least 450 protein groups, at least 500 protein groups, at least 600 protein groups, at least 700 protein groups, at least 800 protein groups, at least 900 protein groups, at least 1000 protein groups, at least 1100 protein groups, at least 1200 protein groups, at least 1300 protein groups, at least 1400 protein groups, at least 1500 protein groups, at least 1600 protein groups, at least 1800 protein groups, at least 2000 protein groups, at least 2500, at least 5000 protein groups, at least 10000 protein groups, at least 15000 protein groups, at least 20000 protein groups, at least 25000 protein groups, at least 30000 protein groups, at least 35000 protein groups, at least 45000 protein groups, at least 50000 protein groups, at least 60000 protein groups, at least 70000 protein groups, at least 80000 protein groups, at least 90000 protein groups, or at least 100000 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising at most 5 protein groups, at most 10 protein groups, at most 20 protein groups, at most 30 protein groups, at most 40 protein groups, at most 50 protein groups, atWSGR Docket No. 53344-800.601most 60 protein groups, at most 80 protein groups, at most 100 protein groups, at most 150 protein groups, at most 200 protein groups, at most 250 protein groups, at most 300 protein groups, at most 400 protein groups, at most 500 protein groups, at most 600 protein groups, at most 800 protein groups, at most 1000 protein groups, at most 1200 protein groups, at most 1500 protein groups, at most 1800 protein groups, at most 2000 protein groups, at most 2500 protein groups, at most 3000 protein groups, at most 4000 protein groups, at most 5000 protein groups, at most 7500 protein groups, at most 10000 protein groups, at most 15000 protein groups, at most 20000 protein groups, at most 25000 protein groups, at most 50000 protein groups, at most 75000 protein groups, or at most 100000 protein groups. A particle disclosed herein can be incubated with a biological sample to form a protein corona comprising from 5 to 2500 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 5 to 50 protein groups. A particle disclosed herein can be incubated with a biological sample to form a protein corona comprising from 10 to 100 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 20 to 100 protein groups. A particle disclosed herein can be incubated with a biological sample to form a protein corona comprising from 20 to 400 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 50 to 500 protein groups. A particle disclosed herein can be incubated with a biological sample to form a protein corona comprising from 100 to 800 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 200 to 1000 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 300 to 1200 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 400 to 1500 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 500 to 2000 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 800 to 2500 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 1000 to 3000 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 1000 to 5000 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 2000 to 10000 protein groups. A particle or particle panel disclosed herein can be incubated with a biological sample to form a protein corona comprising from 5000 to 25000 protein groups. In some cases, several different types of particles can be used, separately or in combination, to identify large numbers of proteins in a particular biological sample. In other words, particles can be multiplexed in order to bind and identify large numbers of proteins in a biological sample. ProteinWSGR Docket No. 53344-800.601corona analysis may compress the dynamic range of the analysis compared to a protein analysis of the original sample.

[0110] A biological sample (e.g., human plasma) comprising a plurality of biomolecules may be contacted to a plurality of particles. The sample may be treated, diluted, or split into a plurality of fractions prior to analysis. For example, a whole blood sample may be fractionated into plasma and erythrocyte portions. Upon contact with the particles, a subset or the entirety of the plurality of biomolecules may adsorb to the particles, thereby forming biomolecule coronas bound to the surfaces of the particles. Unbound biomolecules may be separated from the biomolecule coronas (e.g., through wash steps). The biomolecule coronas, or subsets thereof, may be collected from the particles. Alternatively, biomolecules of the biomolecule coronas may be fragmented or chemically treated while bound to the particles. In some assays, biomolecules (e.g., proteins) are fragmented (e.g., digested) while disposed in the biomolecule coronas to yield biomolecule (e.g., peptide) fragments. Biomolecules (or their chemically treated or fragmented derivatives) may then be analyzed by mass spectrometry to yield data representative of biomolecules from the biological sample. The data may be analyzed to identify a biological state of the biological sample.[OHl] In some embodiments, an example of a biomolecule corona (e.g., protein corona) analysis workflow of the present disclosure which includes: particle incubation with a biological sample (e.g., plasma) under conditions suitable for adsorption of biomolecules from the biological sample to the particles to form biomolecule coronas; partitioning of the particle-plasma sample mixture into a plurality of partitions (e.g., wells on a multi-well plate); particle collection (e.g., with a magnet); a wash step or plurality of wash steps to remove analytes not adsorbed to the particles; resuspension of the particles and the biomolecules adsorbed thereto; biomolecule corona digestion or chemical treatment (e.g., protein reduction and digestion); and analysis of the biomolecule coronas or of biomolecules derived therefrom (e.g., by liquid chromatography-mass spectrometry (LC-MS) analysis). While this example provides parallel analyses across multiple wells of a multi-well plate, a method may comprise a single sample volume or a plurality of sample volumes, for example 2 volumes, 3 volumes, 4 volumes, 5 volumes, 6 volumes, 7 volumes, 8 volumes, 9 volumes, 10 volumes, 11 volumes, 12 volumes, 15 volumes, 18 volumes, 20 volumes, 22 volumes, 24 volumes, 25 volumes, 28 volumes, 30 volumes, 36 volumes, 40 volumes, 48 volumes, 50 volumes, 60 volumes, 70 volumes, 80 volumes, 90 volumes, 96 volumes, 128 volumes, 150 volumes, 192 volumes, 200 volumes, 250 volumes, 256 volumes, 300 volumes, 384 volumes, 400 volumes, 500 volumes, 512 volumes, 600 volumes, or more. For example, the method may be performed on a 96, 192, or 384 well plate. Furthermore, while this example provides contacting a sample with particles prior to partitioning, a method may alternatively comprise partitioning a sample (e.g., into separate wells of a well plate) prior to contacting with particles. Each sample volume may be separately mixed with particles prior to, concurrent with, or subsequent to addition into a partition. In particularWSGR Docket No. 53344-800.601cases, the particles are present in a partition (for example in dry form or in solution) prior to addition of the sample into the partition. In some cases, sample may be added to partitions comprising particles. For example, a well plate may be provided with particles, buffer, and reagents in dry form, such that a method of use may comprise adding solution to the wells to resuspend the particles and dissolve the buffer and reagents, and then adding sample to the wells.

[0112] Protein corona analysis may comprise an automated component. For example, an automated instrument may contact a sample with a particle or particle panel, identify proteins on the particle or particle panel (e.g., digest the proteins on the particle or particle panel and perform mass spectrometric analysis), and generate data for identifying a specific biomolecule or a biological state of a sample. The automated instrument may divide a sample into a plurality of volumes, and perform analysis on each volume or a subset of the plurality. The automated instrument may analyze multiple separate samples, for example by disposing multiple samples within multiple wells in a well plate, and performing parallel analysis on each sample or a subset of samples within the well plate.

[0113] The methods disclosed herein include isolating one or more particle types from a sample or from more than one sample (e.g., a biological sample or a serially interrogated sample). The particle types can be isolated or separated from the sample using a magnet. Moreover, multiple samples that are spatially isolated can be processed in parallel. Thus, the methods disclosed herein provide for isolating or separating a particle type from unbound protein in a sample. A particle type may be separated using methods including but not limited to magnetic separation, centrifugation, fdtration, or gravitational separation. Particle panels may be incubated with a plurality of spatially isolated samples, wherein each spatially isolated sample is in a well in a well plate (e.g., a 96- well plate, a 192-well plate, or a 384-well plate). After incubation, the particle types in each of the wells of the well plate can be separated from unbound protein present in the spatially isolated samples by placing the entire plate on a magnet. This pulls down the superparamagnetic particles in the particle panel. The supernatant in each sample can be removed to remove the unbound protein. These steps (incubate, pull down) can be repeated to effectively wash the particles, thus removing residual background unbound protein that may be present in a sample. This is one example, but one of skill in the art could envision numerous other scenarios in which superparamagnetic particles are rapidly isolated from one or more than one spatially isolated samples at the same time.

[0114] An assay may comprise protein collection of particles, protein digestion, and mass spectrometric analysis (e.g., MS, LC-MS, LC-MS / MS). The digestion may comprise chemical digestion, such as by cyanogen bromide or 2-Nitro-5-thiocyanatobenzoic acid (NTCB). The digestion may comprise enzymatic digestion, such as by trypsin or pepsin. The digestion may comprise enzymatic digestion by a plurality of proteases. The digestion may comprise a protease selected from among the group consisting of trypsin, chymotrypsin, Glu C, Lys C, elastase, subtilisin, proteinase K, thrombin, factor X, Arg C, papaine, Asp N, thermolysine, pepsin, aspartylWSGR Docket No. 53344-800.601protease, cathepsin D, zinc mealloprotease, glycoprotein endopeptidase, proline, aminopeptidase, prenyl protease, caspase, kex2 endoprotease, or any combination thereof. The digestion may cleave peptides at random positions. The digestion may cleave peptides at a specific position (e.g., at methionines) or sequence (e.g., glutamate-histidine-glutamate). The digestion may enable similar proteins to be distinguished. For example, an assay may resolve 8 distinct proteins as a single protein group with a first digestion method, and as 8 separate proteins with distinct signals with a second digestion method. The digestion may generate an average peptide fragment length of 8 to 15 amino acids. The digestion may generate an average peptide fragment length of 12 to 18 amino acids. The digestion may generate an average peptide fragment length of 15 to 25 amino acids. The digestion may generate an average peptide fragment length of 20 to 30 amino acids. The digestion may generate an average peptide fragment length of 30 to 50 amino acids.

[0115] The present disclosure further provides computer systems that are programmed to implement methods of the disclosure. FIG. 28 shows a computer system 2301 that is programmed or otherwise configured to perform any of the methods described herein. The computer system 2301 can regulate various aspects of the present disclosure, such as, for example, searching of data independent acquisition proteomics data from a mass spectrometer to identify detected proteins with reduced global false discovery rates, and / or acquisition of such proteomics data from a large number of samples (e.g., at least 50 samples per day) using one or more autosamplers of a mass spectrometer. The computer system 2301 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.

[0116] The computer system 2301 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 2305, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 2301 also includes memory or memory location 2310 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 2315 (e.g., hard disk), communication interface 2320 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 2325, such as cache, other memory, data storage and / or electronic display adapters. The memory 2310, storage unit 2315, interface 2320 and peripheral devices 2325 are in communication with the CPU 2305 through a communication bus (solid lines), such as a motherboard. The storage unit 2315 can be a data storage unit (or data repository) for storing data. The computer system 2301 can be operatively coupled to a computer network (“network”) 2330 with the aid of the communication interface 2320. The network 2330 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 2330 in some cases is a telecommunication and / or data network. The network 2330 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 2330, in some cases with the aid of the computer systemWSGR Docket No. 53344-800.6012301, can implement a peer-to-peer network, which may enable devices coupled to the computer system 2301 to behave as a client or a server.

[0117] The CPU 2305 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 2310. The instructions can be directed to the CPU 2305, which can subsequently program or otherwise configure the CPU 2305 to implement methods of the present disclosure. Examples of operations performed by the CPU 2305 can include fetch, decode, execute, and writeback.

[0118] The CPU 2305 can be part of a circuit, such as an integrated circuit. One or more other components of the system 2301 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0119] The storage unit 2315 can store files, such as drivers, libraries and saved programs. The storage unit 2315 can store user data, e.g., user preferences and user programs. The computer system 2301 in some cases can include one or more additional data storage units that are external to the computer system 2301, such as located on a remote server that is in communication with the computer system 2301 through an intranet or the Internet.

[0120] The computer system 2301 can communicate with one or more remote computer systems through the network 2330. For instance, the computer system 2301 can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 2301 via the network 2330.

[0121] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 2301, such as, for example, on the memory 2310 or electronic storage unit 2315. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 2305. In some cases, the code can be retrieved from the storage unit 2315 and stored on the memory 2310 for ready access by the processor 2305. In some situations, the electronic storage unit 2315 can be precluded, and machine-executable instructions are stored on memory 2310.

[0122] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.

[0123] Aspects of the systems and methods provided herein, such as the computer system 2301, can be embodied in programming. Various aspects of the technology may be thought of as “products” orWSGR Docket No. 53344-800.601“articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine-readable medium. Machineexecutable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.

[0124] Hence, a machine-readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.WSGR Docket No. 53344-800.601

[0125] The computer system 2301 can include or be in communication with an electronic display 2335 that comprises a user interface (UI) 2340 for providing, for example, a control interface for acquiring mass spectrometry data from one or more samples, a search function for input and identification of peptides and / or proteins from data independent acquisition mass spectrometry data, and / or one or more results outputs providing a user with identified proteins and / or peptides from such data. Examples of UI’s include, without limitation, a graphical user interface (GUI) and webbased user interface.

[0126] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 2305. The algorithm can, for example, facilitate the performance of any of the methods described herein.

[0127] Further provided herein are embodiments according to the invention:

[0128] Embodiment 1. A method, comprising:receiving a plurality of retention times from a mass spectrometer;processing the plurality of retention times using a model configured to calibrate a retention time of the plurality of retention times;determining a first protein structure match based on the processed plurality of retention times; and training a machine learning model to output a second protein structure match, wherein the machine learning model is trained on a dataset comprising the first protein structure match and the plurality of processed retention times.

[0129] Embodiment 2. The method of embodiment 1, wherein receiving the plurality of retention times comprises receiving a data file generated by the mass spectrometer (e.g., a data file corresponding to a single mass spectrometer injection).

[0130] Embodiment 3. The method of any one of embodiments 1 to 2, wherein the mass spectrometer is configured for data independent acquisition (DIA).

[0131] Embodiment 4. The method of any one of embodiments 1 to 3, wherein the received plurality of retention times are stored in a database (e.g., located in a memory of a computational device).

[0132] Embodiment 5. The method of embodiment 4, wherein the database has a R-tree structure (e.g., an original R-tree structure, a packed R-tree structure, or a similar database structure).

[0133] Embodiment 6. The method of any one of embodiments 1 to 5, wherein the first protein structure match comprises an identified peptide.

[0134] Embodiment 7. The method of embodiment 6, wherein the first protein structure match further comprises a corresponding false discovery rate (FDR).

[0135] Embodiment 8. The method of any one of embodiments 1 to 7, wherein the second protein structure match comprises the first protein structure match and an optimized FDR.WSGR Docket No. 53344-800.601

[0136] Embodiment 9. The method of any one of embodiments 1 to 8, wherein the model configured to calibrate the retention time comprises a retention time (RT) alignment model and an RT window filter.

[0137] Embodiment 10. The method of embodiment 9, wherein the RT window filter is derived from the RT alignment model.

[0138] Embodiment 11. The method of any one of embodiments 1 to 10, wherein the model configured to calibrate the retention time comprises a regression model.

[0139] Embodiment 12. The method of any one of embodiments 1 to 11, wherein the model configured to calibrate the retention time comprises a linear discriminant analysis (LDA).

[0140] Embodiment 13. The method of any one of embodiments 1 to 12, wherein the model configured to calibrate the retention time is trained on a dataset comprising empirical retention data and predicted retention data.

[0141] Embodiment 14. The method of embodiment 13, wherein the predicted retention data is obtained from a target library.

[0142] Embodiment 15. The method of embodiment 14, wherein the targets in the target library are peptides.

[0143] Embodiment 16. The method of any one of embodiments 1 to 15, wherein receiving the plurality of retention times further comprises receiving a plurality of mass-to-charge (m / z) values, wherein each retention time of the plurality of retention times is paired with a corresponding m / z value of the plurality of m / z values.

[0144] Embodiment 17. The method of embodiment 16, further comprising:computing a mass calibration curve; andrecalibrating the received plurality of corresponding m / z values using the computed mass calibration curve.

[0145] Embodiment 18. The method of embodiment 17, wherein computing the mass calibration curve and recalibrating the received plurality of corresponding m / z values is performed after calibrating the plurality of retention times.

[0146] Embodiment 19. The method of any one of embodiments 16 to 18, wherein determining the first protein structure match is based on the processed plurality of retention times and the recalibrated plurality of corresponding m / z values.

[0147] Embodiment 20. The method of any one of embodiments 1 to 19, wherein receiving the plurality of retention times further comprises receiving a plurality of signal intensities, wherein each retention time of the plurality of retention times is paired with a corresponding signal intensity of the plurality of signal intensities.WSGR Docket No. 53344-800.601

[0148] Embodiment 21. The method of embodiment 20, wherein determining the first protein structure match is based on the processed plurality of retention times, the plurality of corresponding signal intensities, and, optionally, the recalibrated plurality of corresponding m / z values.

[0149] Embodiment 22. The method of any one of embodiments 1 to 21, wherein receiving the plurality of retention times further comprises receiving a plurality of ion mobility values, wherein each retention time of the plurality of retention times is paired with a corresponding ion mobility value.

[0150] Embodiment 23. The method of embodiment 22, further comprising: computing an ion mobility calibration curve, and recalibrating the received plurality of corresponding ion mobility values using the computed ion mobility calibration curve.

[0151] Embodiment 24. The method of embodiment 23, wherein computing the ion mobility calibration curve and recalibrating the received plurality of corresponding ion mobility values is performed after calibrating the plurality of retention times.

[0152] Embodiment 25. The method of any one of embodiments 22 to 24, wherein determining the first protein structure match is based on the processed plurality of retention times, the recalibrated plurality of corresponding ion mobility values, and, optionally, the recalibrated plurality of corresponding m / z values and / or the plurality of corresponding signal intensities.

[0153] Embodiment 26. The method of any one of embodiments 1 to 25, wherein the machine learning model comprises a neural network.

[0154] Embodiment 27. The method of any one of embodiments 1 to 26, wherein the second protein structure match output by the machine learning model comprises an optimized FDR.

[0155] Embodiment 28. The method of any one of embodiments 1 to 27 further comprising:determining a plurality of first protein structure matches, wherein the machine learning model is trained on the plurality of first protein structure matches, the plurality of processed retention times, and, optionally, the recalibrated plurality of corresponding m / z values and / or the plurality of corresponding signal intensities and / or the recalibrated plurality of corresponding ion mobility values.

[0156] Embodiment 29. The method of embodiment 28, wherein each protein structure match of the plurality of first protein structure matches has an FDR that is below a pre-selected threshold.

[0157] Embodiment 30. The method of embodiment 29, wherein the pre-selected FDR threshold for the first protein structure matches is 50% FDR or less (e.g., 45%, 40%, 35%, 30%, 25%, or less).

[0158] Embodiment 31. The method of any one of embodiments 28 to 30, further comprising: outputting a list comprising a plurality of second protein structure matches, wherein each second protein structure match passes a pre-selected FDR threshold.WSGR Docket No. 53344-800.601

[0159] Embodiment 32. The method of embodiment 31, wherein the pre-selected FDR threshold for the plurality of second protein structure matches is user-selected.

[0160] Embodiment 33. The method of embodiment 31 or 32, wherein the pre-selected FDR threshold for the plurality of second protein structure matches is an FDR of 1% or less.

[0161] Embodiment 34. A method, comprising:receiving, at a plurality of processing units, mass spectrometry data from at least one mass spectrometer;processing the mass spectrometry data, wherein the processing comprises calibrating a retention time of the mass-spectrometry data; anddetermining a plurality of protein structure matches.

[0162] Embodiment 35. The method of embodiment 34, wherein the processing comprises performing the method of any one of embodiments 1 to 33.

[0163] Embodiment 36. The method of embodiment 34 or 35, wherein the plurality of protein structure matches has a false match rate of less than about 0.5%.

[0164] Embodiment 37. The method of any one of embodiments 34 to 36, wherein the false match rate at a precursor level is less than about 1%.

[0165] Embodiment 38. The method of any one of embodiments 34 to 37, wherein the false match rate at a protein group level is less than about 1%.

[0166] Embodiment 39. The method of any one of embodiments 34 to 38, wherein the plurality of processing units comprises at least two nodes.

[0167] Embodiment 40. The method of embodiment 34, wherein the calibrating the retention time comprises calibrating an ion retention time based on an average mass calibration curve.

[0168] Embodiment 41. A method for protein identification and / or quantification by mass spectrometry (MS) experiments, comprising:obtaining a plurality of mass spectra utilizing data-independent acquisition (DIA); processing by linear discriminant analysis (LDA) input data representing complex biological samples to build a model that correlates predicted retention times with empirical values;applying a retention time window filter based on the model generated by LDA; recalibrating mass-to-charge (m / z) values for all ions in the plurality of mass spectra using average mass calibration curves across batches; anddetermining peptide-spectrum matches (PSMs) based at least in part on the model generated by LDA.

[0169] Embodiment 42. The method of embodiment 41, wherein the method further comprises a high-throughput proteomics workflow for processing at least 1000 or more data files (e.g., at least 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10,000,WSGR Docket No. 53344-800.60111,000, 12,000, 13,000, 14,000, 15,000, 20,000, 25,000, or more datafiles) comprising the plurality of mass spectra within a 24-hour period.

[0170] Embodiment 43. The method of any one of embodiments 41 to 42, wherein the method processes a population-scale proteomics study with an FDR of less than about 50% (e.g., less than about 45%, 40%, 35%, or 30%).

[0171] Embodiment 44. The method of any one of embodiments 41 to 42, wherein a false discovery rate of the method is less than about 5% (e.g., less than about 4%, 3%, 2%, 1.5%, or 1%).

[0172] Embodiment 45. The method of any one of embodiments 41 to 42, wherein the LDA processing is performed individually on each MS data file comprising mass spectra of the plurality of mass spectra.

[0173] Embodiment 46. The method of any one of embodiments 41 to 42, wherein the LDA processing comprises a pairwise analysis of peptide target spectra and corresponding decoy spectra.

[0174] Embodiment 47. The method of embodiment 46, wherein the peptide target spectra are a subset of peptide target spectra from a peptide target library.

[0175] Embodiment 48. The method of embodiment 47, wherein the peptide target spectra, the decoy spectra, and / or the peptide target library is / are stored in an accessible memory database having an R-tree structure, for example comprising an original R-tree, a packed R-tree, or a similar database structure.

[0176] Embodiment 49. The method of any one of embodiments 45 to 48, wherein the peptide target spectra and the corresponding decoy spectra are represented by multi-dimensional feature vectors.

[0177] Embodiment 50. The method of embodiment 49, wherein the multi-dimensional feature vectors comprise at least 50 features (e.g., at least 75, 100, 110, 120, 130, 140, 150, or more features).

[0178] Embodiment 51. The method of any one of embodiments 41 to 50, wherein the average mass calibration curves are generated using data output from the LDA processing.

[0179] Embodiment 52. The method of any one of embodiments 41 to 51, wherein determining the PSMs comprises performing regression analysis.

[0180] Embodiment 53. The method of any one of embodiments 41 to 52, wherein determining the PSMs comprises performing LDA.

[0181] Embodiment 54. The method of any one of embodiments 41 to 53, wherein determining the PSMs comprises generating feature vectors representative of each peptide target spectrum of a peptide target library, and scoring the feature vectors against mass spectra of the plurality of mass spectra having recalibrated m / z values.WSGR Docket No. 53344-800.601

[0182] Embodiment 55. The method of embodiment 54, wherein each feature vector comprises at least 100 features (e.g., at least 120, 130, 140, 150 features; or about 160 to about 200 features, about 170 to about 190 features, or about 180 features).

[0183] Embodiment 56. The method of any one of embodiments 41 to 55, wherein determining the PSMs is further based at least in part on the retention time window filter and / or the recalibrated m / z values.

[0184] Embodiment 57. The method of any one of embodiments 41 to 56 further comprising normalization of ion counts or intensities prior to the step of determining PSMs.

[0185] Embodiment 58. The method of any one of embodiments 41 to 57 further comprising calculating precursor quantities for the determined PSMs.

[0186] Embodiment 59. The method of embodiment 58, wherein calculating a precursor quantity comprises summing detected fragment intensities.

[0187] Embodiment 60. The method of embodiment 59, wherein summing the detected fragment intensities comprises summing at least the detected fragment intensities over the highest single scan.

[0188] Embodiment 61. The method of embodiment 59 or 60, wherein summing the detected fragment intensities comprises summing the detected fragment intensities over all scans.

[0189] Embodiment 62. A method for searching mass spectrometry data, the method comprising:obtaining a mass spectrometry (MS) data file comprising retention times and corresponding mass-to-charge (m / z) values;performing a first search of the MS data file using a first plurality of peptide target spectra from a library of peptide targets, thereby identifying a first set of peptide-spectrum matches (PSMs);calculating an optimized retention time (RT) window filter using predicted retention times for target peptides represented in the first set of PSMs and their corresponding observed retention times in the MS data file;computing a mass calibration curve using predicted m / z values for the target peptides in the first set of PSMs and corresponding observed m / z values in the MS data file;recalibrating the m / z values in the MS data file using the computed mass calibration curve, thereby generating a recalibrated MS data file;calculating optimized mass tolerances by performing a series of iterative searches on the recalibrated MS data file, using a second plurality of peptide target spectra from the peptide target library;scoring the peptide target library against the recalibrated MS data file using the optimized RT window filter and the optimized mass tolerances, thereby generating a set of optimized PSMs;WSGR Docket No. 53344-800.601calculating optimized false discovery rates (FDRs) for the optimized PSMs in the set of optimized PSMs using a machine learning algorithm.

[0190] Embodiment 63. The method of embodiment 62, further comprising outputting a list of optimized PSMs having optimized FDRs below a pre-selected FDR threshold.

[0191] Embodiment 64. The method of embodiment 63, wherein the pre-selected FDR threshold is user-selected.

[0192] Embodiment 65. The method of embodiment 63 or 64, wherein the pre-selected FDR threshold is an FDR of 1% (e.g., 2.0%, 1.5%, 1%, or less).

[0193] Embodiment 66. The method of any one of embodiments 63 to 65, wherein the list of optimized PSMs comprises a feature vector and related meta-data for each optimized PSM in the list.

[0194] Embodiment 67. The method of embodiment 66, wherein the meta-data includes a score (e.g., single-value score) for the feature vector, a sequence for the matched peptide, a retention time for the matched peptide, optionally ion mobility time for the matched peptide, etc.)

[0195] Embodiment 68. The method of any one of embodiments 62 to 67, wherein the MS data file corresponds to a single MS injection.

[0196] Embodiment 69. The method of any one of embodiments 62 to 68, wherein the MS data file has an mzML or Parquet file format.

[0197] Embodiment 70. The method of any one of embodiments 62 to 69, wherein the first plurality of peptide target spectra corresponds to about 0.1% to about 25% of the peptide targets in the peptide target library (e.g., up to about 20%, 15%, 10%, 7%, 5%, 4%, 3%, 2%, or 1% of the peptide targets in the peptide target library).

[0198] Embodiment 71. The method of any one of embodiments 62 to 70, wherein the first plurality of peptide target spectra comprises about 20,000 peptide target spectra (e.g., at least about 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000 or more peptide target spectra).

[0199] Embodiment 72. The method of any one of embodiments 62 to 71 further comprising:obtaining the peptide target library; andstoring the peptide target library in a memory database having an R-tree structure.

[0200] Embodiment 73. The method of embodiment 72, wherein the R-tree structure is an original R-tree, a packed R-tree, or has a similar database structure.

[0201] Embodiment 74. The method of any one of embodiments 62 to 73, wherein each peptide target spectrum in the first plurality of peptide target spectra is represented as a vector having a plurality of features.

[0202] Embodiment 75. The method of embodiment 74, wherein each vector representing a peptide target spectrum comprises at least 50 features (e.g., at least 75, 100, 110, 120, 130, 140, 150, or more features).

[0203] Embodiment 76. The method of any one of embodiments 62 to 75 further comprising:WSGR Docket No. 53344-800.601prior to building an RT alignment model, generating a decoy spectrum corresponding to each peptide target spectrum in the first plurality of peptide target spectra.

[0204] Embodiment 77. The method of embodiment 76, wherein generating a decoy spectrum corresponding to a target spectrum comprises determining a spectrum corresponding to a mutated peptide sequence derived from a peptide sequence corresponding to the peptide target spectrum.

[0205] Embodiment 78. The method of embodiment 76 or 77, wherein decoy spectrum is represented as a vector having a plurality of features.

[0206] Embodiment 79. The method of embodiment 78, wherein each decoy vector representing a peptide target spectrum comprises at least 50 features (e.g., at least 75, 100, 110, 120, 130, 140, 150, or more features).

[0207] Embodiment 80. The method of any one of embodiments 76 to 79, wherein building the RT alignment model comprises scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file.

[0208] Embodiment 81. The method of embodiment 80, wherein scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file comprises regression analysis (e.g., logistic regression).

[0209] Embodiment 82. The method of embodiment 80, wherein scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file comprises linear determinant analysis (LDA).

[0210] Embodiment 83. The method of any one of embodiments 80 to 82, wherein scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file comprises, for each potential PSM, generating a score for each of the peptide target spectrum and the corresponding decoy spectrum and determining, based upon a difference between those scores, whether the peptide target spectrum has a PSM with the MS data file.

[0211] Embodiment 84. The method of any one of embodiments 62 to 83, wherein building the RT alignment model is an iterative process.

[0212] Embodiment 85. The method of embodiment 84, wherein building the RT alignment model comprises using a first plurality of target spectra from the target library in combination with the MS data file to build a first RT alignment model, using a second plurality of target spectra from the target library in combination with the MS data file to build a second RT alignment model, and evaluating whether the first RT alignment model and the second RT alignment model are sufficiently convergent.

[0213] Embodiment 86. The method of any one of embodiments 62 to 85, wherein calculating the optimized RT window comprises using splines generated from a subset of PSMs in the first set ofPSMs.WSGR Docket No. 53344-800.601

[0214] Embodiment 87. The method of embodiment 86, wherein each PSM in the subset of PSMs meets a pre-selected false discovery rate (FDR) (e.g., 10%, 8%, 6%, 5%, 4%, 3% or less).

[0215] Embodiment 88. The method of embodiment 86 or 87, wherein the subset of PSMs comprises at least 200 PSMs (e.g., at least 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 2000, 2500, 3000, or more PSMs).

[0216] Embodiment 89. The method of any one of embodiments 62 to 88, wherein the mass calibration curve is computed using linear regression.

[0217] Embodiment 90. The method of embodiment 89, wherein the mass calibration curve is computed using a least squares regression to fit a polynomial.

[0218] Embodiment 91. The method of embodiment 90, wherein the polynomial is a second order polynomial.

[0219] Embodiment 92. The method of any one of embodiments 62 to 91, wherein the searches in the series of searches of the recalibrated MS data file use an ascending set of values for mass tolerance.

[0220] Embodiment 93. The method of any one of embodiments 62 to 92, wherein a first search in the series of searches of the recalibrated MS data file uses a mass tolerance of 3 ppm.

[0221] Embodiment 94. The method of embodiment 93, wherein subsequent searches in the series of searches use mass tolerances of 4 ppm, 5ppm, 6 ppm, etc., up to 50 ppm.

[0222] Embodiment 95. The method of any one of embodiments 62 to 94, wherein the series of searches of the recalibrated MS data file is stopped when PSMs resulting from the series of searches reaches a maximum (e.g., a local maximum; or when the number of target counts does not increase after one, two, three, or more further searches).

[0223] Embodiment 96. The method of any one of embodiments 62 to 95, wherein scoring the target library against the recalibrated MS data file comprises performing regression analysis.

[0224] Embodiment 97. The method of any one of embodiments 62 to 96, wherein scoring the target library against the recalibrated MS data file comprises performing LDA.

[0225] Embodiment 98. The method of any one of embodiments 62 to 97, wherein scoring the target library against the recalibrated MS data file comprises generating feature vectors representative of each peptide target spectrum from the peptide target library, and scoring the feature vectors against the recalibrated MS data file.

[0226] Embodiment 99. The method of embodiment 98, wherein each feature vector comprises at least 100 features (e.g., at least 120, 130, 140, 150 features; or about 160 to about 200 features, about 170 to about 190 features, or about 180 features).

[0227] Embodiment 100. The method of any one of embodiments 62 to 99, wherein scoring the target library against the recalibrated MS data file further comprising calculating precursor quantities for the set of optimized PSMs.WSGR Docket No. 53344-800.601

[0228] Embodiment 101. The method of embodiment 100, wherein calculating a precursor quantity comprises summing detected fragment intensities.

[0229] Embodiment 102. The method of embodiment 101, wherein summing the detected fragment intensities comprises summing at least the detected fragment intensities over the highest single scan.

[0230] Embodiment 103. The method of embodiment 101 or 102, wherein summing the detected fragment intensities comprises summing the detected fragment intensities over all scans.

[0231] Embodiment 104. The method of any one of embodiments 62 to 103, wherein the machine learning algorithm is a neural network.

[0232] Embodiment 105. The method of embodiment 104, wherein the neural network comprises a classifier having three fully-connected layers and a configurable hidden layer width using ReLU activation.

[0233] Embodiment 106. The method of embodiment 104 or 105, wherein the neural network is trained using binary cross-entropy loss and multi-fold cross-validation.

[0234] Embodiment 107. The method of any one of embodiments 62 to 106 further comprising:obtaining a plurality of MS data files, each comprising retention times and corresponding m / z values; andsearching each of the MS data files in the plurality of MS data files.

[0235] Embodiment 108. The method of embodiment 107, wherein the plurality of MS data files comprises at least 100 MS data files (e.g., at least 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 25000, or more data files).

[0236] Embodiment 109. The method of embodiment 107 or 108, wherein each MS data file in the plurality of MS data files is stored and / or searched in a single node (e.g., a single computer or a single processing unit in a distributed computing environment).

[0237] Embodiment 110. The method of any one of embodiments 107 to 109, wherein each MS data file in the plurality of MS data files is searched independently and / or wherein the MS data files in the plurality of MS data files are searched in parallel.

[0238] Embodiment 111. A method for analyzing a mass spectrometry data file comprising: processing a mass spectrometry data file using a library of predicted peptide spectra to identify i) a plurality of protein structure matches; ii) a first set of scores associes with the plurality of protein structure matches; and iii) a second set of scores associated with a plurality of decoy matches;dividing the first set of scores and the second set of scores into a plurality of blocks, each block containing a subset of the protein structure matches;WSGR Docket No. 53344-800.601computing, for each block, a count of the number of protein structure matches and a count of the number of decoy matches within that block; andusing the counts of the number of protein structure matches and the number of decoy matches within each block to estimate a false discovery rate (FDR) for each block; and aggregating the FDR estimates across the plurality of blocks to yield an overall FDR estimate for the plurality of protein structure matches.

[0239] Embodiment 112. The method of embodiment 111, wherein the mass spectrometry data file is generated from a biological sample.

[0240] Embodiment 113. The method of embodiment 112, wherein the method comprises calculating, for each block, an upper-bound q-value estimate and a lower-bound q-value estimate based on the first set of scores and the second set of scores within that block.

[0241] Embodiment 114. The method of any one of embodiments 111 to 113, wherein the computing is performed using a distributed computing environment comprising a plurality of nodes.

[0242] Embodiment 115. The method of any one of embodiments 111 to 114, wherein the number of blocks is between 20,000 and 100,000.

[0243] Embodiment 116. The method embodiment 115, wherein the number of blocks is between 30,000 and 80,000.

[0244] Embodiment 117. The method of any one of embodiments 111 to 116, wherein the method further comprises selecting an acceptable relative error for quantization.

[0245] Embodiment 118. The method of embodiment 117, wherein the acceptable relative error for quantization is less than 0.0005, 0.0004, 0.0003, 0.0002, or 0.0001.

[0246] Embodiment 119. A non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to perform operations comprising a method of any one of embodiments 1 to 118.

[0247] Embodiment 120. A system for performing quantitative DIA proteomics, comprising:a computer configured to perform the methods of any one of embodiments 1 to 118;a data storage system for managing and processing one or more mzML files, each independently comprising data-independent analysis mass spectra of one or more biological samples; anda distributed computing environment (e.g., Apache Spark) for scalable execution of the data analysis pipeline.

[0248] Embodiment 121. The system of embodiment 120, wherein the distributed computing environment is scalable to analyze at least 100 samples per day (e.g., at least 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, or more samples per day).

[0249] Embodiment 122. The system of embodiment 120 or 121, having the configuration shown in FIG. 2.WSGR Docket No. 53344-800.601

[0250] Embodiment 123. The system of embodiment 120, 121, or 122, having the configuration shown in FIG. 3.ExamplesExample 1 : Implementation of Methods described herein providing a Fast, Sensitive, and Accurate Search Engine for Quantitative DIA Proteomics

[0251] Large-scale mass spectrometry (MS) experiments enabled comprehensive identification and quantification of proteins in complex biological samples. However, challenges were presented by datasets containing thousands of injections in maintaining sensitivity, quantitative accuracy, and a well-controlled false discovery rate (FDR). Traditional pipelines often struggled to scale while preserving statistical rigor and confidence in results.

[0252] An example embodiment of a proteomics search engine provided herein was introduced as an efficient and highly scalable data-independent acquisition (DIA) MS search engine designed to ensure robust identification and quantification while maintaining stringent FDR control across large-scale datasets. Advanced algorithms optimized for sensitivity and reproducibility were utilized by an example embodiment of a proteomics search engine provided herein, achieving consistent performance even with extensive and diverse input data.

[0253] The ability of the example embodiment of a proteomics search engine provided herein to deliver accurate and reliable results in large-scale proteomics experiments was demonstrated, balancing sensitivity and specificity for high-confidence protein quantification.

[0254] Runtime performance was maximized, and computing costs were minimized by the example embodiment of a proteomics search engine provided herein, starting from accelerated parallel file loading, followed by efficient data structures and algorithms for searching Peptide-Spectrum Matches (PSMs). Small library batches were processed using Linear Discriminant Analysis (LDA) to build a model correlating predicted retention time with empirical values, setting a retention time window filter. Average mass calibration curves across batches were used to recalibrate every ion in the data file.

[0255] Mass tolerances for extracting transition traces were optimized by iteratively processing small library batches at values from 3 to 50 ppm. The entire library was processed using LDA, with resulting PSMs used to train a neural network classifier. PSMs below 50% FDR were reported for downstream processing with a scalable Spark-based pipeline.

[0256] Thirty-seven plasma samples were analyzed with the Proteograph® XT assay, resulting in 74 DIA injections on Orbitrap™ Astral™. Candidate PSMs were generated using an example embodiment of a proteomics search engine provided herein and DIA -NN 1.8.1 (in “library-free” mode). Both analyses employed the same library of 8,621 proteins detectable in human plasma, combined with 8,730 Arabidopsis and 116 contaminant proteins. The AlphaPeptDeep model for the Orbitrap mass analyzer was used for spectrum prediction. Individual injections were processed inWSGR Docket No. 53344-800.601parallel, utilizing 8 CPUs for each file. Both sets of results were filtered to 1% PSM and precursor FDR. Proteins were grouped parsimoniously, and protein group (PG) results were further filtered to 1% PG FDR.

[0257] A total of 1,714,526 PSMs were produced by DIA-NN from 73,695 precursors and 4,303 PGs, estimating PSM-level FDR locally per injection. Parallel processing of individual injections, followed by global processing on a single node (32 CPUs), took 162 minutes.

[0258] The example embodiment of a proteomics search engine provided herein produced 1,692,289 PSMs from 71,424 precursors and 4,234 PGs, estimating PSM-level FDR globally.Parallel processing of individual injections, followed by global processing using an autoscaling Spark cluster (2-8 nodes, 16-64 total CPUs), took 39 minutes.

[0259] The accuracy of each FDR method was assessed by calculating False Match Rates (FMR) using the entrapment approach. At the PSM level, both analyses had acceptable FMR around 0.4%. FMR was also correctly controlled by the example embodiment of a proteomics search engine provided herein at the precursor and PG levels, with rates of 0.95% and 0.97% respectively.However, DIA-NN exhibited inflated FMR of 1.63% at the precursor level and 1.33% at the PG level.

[0260] The same precursors were largely identified by both methods, with 94% of the precursors identified by the example embodiment of a proteomics search engine provided herein also identified by DIA-NN. A slight reduction in PG sparsity in the results of the example embodiment of a proteomics search engine provided herein was also observed, along with a 3x reduction in singlepeptide PGs (“one-hit wonders”).

[0261] Preliminary analysis showed that sensitivity comparable to existing tools was exhibited by the example embodiment of a proteomics search engine provided herein, with improved performance and FDR estimation.Example 2: Coupling of an example implementation of a PSE to an MPF to provide an informatic pipeline for population scale proteomics

[0262] To complement scoring from the example embodiment of a proteomics search engine provided herein, the MPF pipeline was developed as a scalable pipeline to enable flexible, accurate, and transparent FDR control, which further implements methods described herein. The MPF pipeline was implemented in the Apache Spark framework to enable high-speed and low-resource application for studies at scale.

[0263] As noted in Example 1, the example embodiment of a proteomics search engine provided herein is a peptide-centric DIA search engine that enables high accuracy, high sensitivity measurements. For the example embodiment of a proteomics search engine provided herein integration in cloud-scalable pipelines with MPF enables population scale observations. ComparedWSGR Docket No. 53344-800.601to existing algorithms, the example embodiment of a proteomics search engine provided herein enables faster search with greater accuracy in MMCC ground-truth experiment frameworks.

[0264] The example embodiment of a proteomics search engine provided herein is a peptide-centric DIA search engine that processes community standard mzML files (though other file types could be utilized). It is highly optimized, with multi-threading, optimized I / O and processing implementations for low CPU and RAM usage and extremely fast overall processing speed on Intel and ARM CPUs. Small batches of a test library were scored with linear discriminant analysis (LDA) to build a model correlating predicted retention time with empirical values. An optimal retention time window filter was established for the search. Next, recalibration of m / z was performed for all ions, and an optimal mass tolerance was determined. Finally, the entire library was processed using LDA, with resulting peptide-spectrum matches (PSMs) used to train a neural network classifier.

[0265] The MPF pipeline framework used fast distributed implementations for global FDR estimation, protein inference, and quantitative roll-up. MPF is highly flexible, with plug-in modules for each step. The entire MPF workflow was implemented in Python using Apache Spark for high performance and scalability in cluster or cloud computing environments or running locally.

[0266] Comparative analysis was done with fair FDR filtering against DIA -NN vl.8.1 (Q. Value, Global / Lib.Q. Value, and Global / Lib.PG.Q. Value in DIA -NN’s report). Matrix Matched Calibration Curve (MMCC) analysis was also executed in EncyclopeDIA v3 with appropriate filters for MMCC ion-centric analysis filtered at Global / Lib q- value < 0.01.Example 3 : Benchmarking performance of the example embodiment of a proteomics search engine provided herein on diverse sample sets

[0267] Evaluation of the example embodiment of a proteomics search engine provided herein was performed on diverse datasets in a single-node environment executing an example MPF provided herein. Three publicly available datasets were downloaded and searched on equivalent 32 AMD CPU, 64 GB RAM computers, in a single Docker container to demonstrate performance without cloud scalability relative to conventional search of DIA -NN vl.8.1. Each algorithm was used to search: Astral Cell Lysate (“Hour Proteome”) with narrow DIA windows, HeLa lysate on a Q-Exactive HF with wide DIA windows, and an eye lens tissue measured on an Exploris 480 with wide DIA windows. The eye lens data was searched with inclusion of one variable deamidation in the database. For the Hour Proteome data, a varying proportion of decoy sequences was added to the library to demonstrate the effects of library size tradeoff on search speed and sensitivity, and to assess accuracy of estimated q-values. Identifications on the Precursor level for each dataset (n = 3) are shown in FIG. 4. Search time for each respective dataset is illustrated in FIG. 5.

[0268] Evaluation of the search performance of the example embodiment of a proteomics search engine provided herein in a cloud computing environment was also performed. The MPF pipelineWSGR Docket No. 53344-800.601architecture is well-suited to cloud environments, using parallel batch processing for execution of the proteomics search engine, along with scalable distributed implementations for other steps.Performance parameters for the search engine, including cost per file and search time, are shown in FIGs. 6A and 6B.

[0269] FIG. 7 illustrates end-to-end execution time for an example embodiment of a proteomics search engine and MPF provided herein compared with a cloud-based scalable DIA-NN pipeline. Unlike DIA-NN, the example embodiment of a proteomics search engine can utilize ARM processors for additional cost reduction, as shown in FIG. 8.

[0270] FIG. 9 illustrates an extended evaluation of the “hour proteome” dataset, including identifications for each respective search using an example PSE provided herein. Three different gradient times for the Orbitrap™ Astral™ were tested, 15 minutes, 30 minutes, and 60 minutes. Only results from the target library were used. FIG. 10 illustrates precision analysis on raw interpreted intensity for all precursors and shared precursors in the datasets of FIG. 9. DIA-NN was operated in "High Precision, Robust LC" mode.

[0271] FIG. 11 illustrates Venn diagrams of precursors measured in all replicates for respective analysis sets of the “hour proteome” dataset of FIG. 9, using example embodiments of PSEs provided herein. FIG. 12 illustrates alignment of intensities for PSMs shared between searches in the FIG. 9 datasets. 10% of PSMs are represented, selected at random, including alignment of observed retention times and residual difference.

[0272] FIG. 13 illustrates protein group analysis from an Astral cohort study showing a histogram of peptides per protein group. FIG. 14 illustrates protein group completeness in the 37-sample study of Example 1. FIG. 15 illustrates precursor entrapment with an Arabadopsis or Shuffle Sequence FASTA concatenation.

[0273] FIGs. 16A and 16B illustrate accuracy assessment with Matrix Matched Calibration Curves using example ground truth data to assess search algorithm figures of merit, including: limits of detection and limits of quantification for an example embodiment of a proteomics search engine provided herein search, DIA-NN vl.8.1, and EncyclopeDIA v3. Analysis was done on a dataset of Yeast peptides spiked into N15 labeled Yeast peptides. Filtering was performed on a global level for each method (Precursor measurements for an example embodiment of a proteomics search engine provided herein and DIA-NN, Peptide for EncyclopeDIA). DIA-NN was also filtered at local level, consistent with developer recommendations. The first row of FIGs. 16A and 16B illustrate figure of merit quantile distributions relative to total entries at 1% FDR at specified LOD / Q thresholds for alternate embodiments of PSEs provided herein compared to other methods. The second row of FIGs. 16A and 16B illustrate Precursor (or Peptide for EncyclopeDIA) count at specified LOD / Q thresholds. The third row of FIGs. 16A and 16B illustrate count distributions at specified LOD / QWSGR Docket No. 53344-800.601thresholds for shared sequences between methods. When more than one charge state was observed, min LOD / Q was used.Example 4: Eye Lens Case Study

[0274] The eye lens is a highly specialized, heterogeneous tissue organized in concentric layers of fiber cells formed throughout life. As lens fiber cells mature, they lose their organelles and persist indefinitely, accumulating age-related PTMs over decades of aging including pervasive deamidation. The small mass shift associated with deamidation (+0.984016 Da) poses substantial challenge for DIA identification algorithms as it competes against the nearby isotopic peak, making this dataset an ideal stress test for evaluating non-labile PTM handling. A dataset of three lens regions - Cortex, Outer Nucleus and Inner Nucleus - was analyzed, and identification performance of an example embodiment of a PSE disclosed herein was compared across modified and unmodified peptides (FIG. 17). PSE identified more precursors than DIA -NN in both library-free and MBR workflows, and detected a larger number of deamidated peptides, particularly in MBR searches, recapitulating the expected biological gradient from cortex to nucleus. Although DIA-NN reports more protein groups, this difference likely reflects divergent strategies for protein inference, including treatment of single-peptide protein groups.

[0275] FIG. 18 illustrates a comparison of precursor-level CV between methods using the same dataset as FIG. 17. FIGs. 19A and 19B illustrate alignment of observed intensities between the PSE and DIA-NN, 10% of PSMs selected at random. FIG. 19C illustrates comparison of library dot products for all identifications using the same dataset as FIG. 17. Skyline independently observed dot products for library free and MBR searches. Retention times observed in respective report files were used for alignment. FIG. 19D shows alignment of precursor intensities between unique and common PSM data subsets of the FIG. 18 dataset between the PSE and DIA-NN.

[0276] FIG. 20 illustrates an Empirical Cumulative Distribution Function (ECDF) plot of library dot products for all identifications using the same dataset as FIG. 17. Skyline independently observed dot products for library free and MBR searches. The library free search utilized the same predicted library for each search. MBR empirical libraries were used for MBR spectral dot product. Retention times observed in respective report files were used for alignment. FIG. 21 illustrates an ECDF plot similar to that shown in FIG. 20 but focused on only deamidated PSMs.

[0277] To compare match quality, chromatographic traces were extracted in Skyline and computed dot products between observed and library entries (FIG. 22). Across all lens regions, PSE yielded a higher median dot product for precursors identified by both engines (-0.7 -2.1% improvement), indicating stronger spectral and retention-time support for its identifications. Differences were more pronounced for engine-unique precursors: in library-free mode, PSE’s unique matches exhibited 4-7% higher absolute median dot products than DIA-NN’ s, and in MBR mode this gap widenedWSGR Docket No. 53344-800.601substantially, with a 22-23% difference in medians. These results demonstrate that PSE maintains match quality during evidence transfer, whereas DIA-NN’s transferred identifications exhibited deteriorated chromatographic support. Trends remained consistent when restricting analysis to deamidated precursors, highlighting robust PTM handling in this challenging dataset.

[0278] The quantitative evidence further supports these observations. Although PSE reports slightly higher precursor intensities overall - with unique identifications generally less intense for both tools - DIA-NN’s unique identifications tend to fall at very low intensities, suggesting increased susceptibility to lower-confidence matches to noise or low-intensity peaks. Residual retention time analysis corroborates this: although most precursors align well between engines, deamidated species exhibit substantially larger RT residuals in DIA-NN (up to ~10 min on a 90-min gradient), consistent with less stable localization of match or misassignment. This effect may be exacerbated by common structural isomers found in the lens that exhibit non-systematic RT shifts. Together, these results indicate that PSE provides higher-quality identifications, leading to more stable quantification of both modified and unmodified peptides, reinforcing its strengths in complex, PTM-rich biological systems.Example 5 : Cancer Case Study

[0279] To evaluate performance in a typical label-free discovery study, a plasma dataset containing 20 cancer and 20 healthy subjects prepared with the Seer Proteograph® XT assay and acquired on an Orbitrap™ Astral™ MS (80 total injections) was re-analyzed. Data were processed using the non-redundant SwissProt library as well as 100% and 200% expansion libraries.

[0280] All PSE runs were completed in one hour in the cloud environment, whereas 45 hours were required by DIA-NN to process six analyses sequentially on an c7a,16xlarge AWS instance.Identification performance at matched 1% PSM-, precursor-, and protein-group FDR thresholds is summarized in FIG. 10. Precursor counts were comparable between engines (FIG. 19) with more single-peptide protein groups being reported by DIA-NN, accounting for 6-9% of its reported identifications.

[0281] Run-to-run quantitative stability was next compared. Distributions of unnormalized precursor CVs were similar between engines (FIGs. 25A-C), reflecting the substantial biological heterogeneity among plasma samples. Increased CV under MBR was shown by both pipelines, but a larger shift towards high-CV precursors was displayed by DIA-NN, suggesting that more variable or lower-confidence matches are included in its transferred identifications.

[0282] To assess discovery performance, differential-abundance (DA) analysis was conducted across all precursor-nanoparticle pairs, controlling the DA false discovery rate at 1% (FIGs. 25A-C). Most significant precursors were shared between engines, with additional candidates not detected by the other being identified by each tool. Unique discoveries were often associated with higherWSGR Docket No. 53344-800.601retention-time variability or detections near the column void volume - features that challenge both the assignment and quantitative stability. These effects were more frequent in DIA-NN, consistent with its higher CV distributions, and may result in increased erroneous DA feature discovery.

[0283] The reduction in significant precursors under MBR for both pipelines likely reflects more complete measurement of variable signals. In contrast, library -free workflows are more susceptible to left-censoring of low-abundance precursors, which can artificially lower variance and inflate significance. Together, these results illustrate that robust quantitative behavior in a heterogenous biological cohort is maintained by PSE while computation efficiencies advantages for large-scale discovery studies are provided.Example 6: Astral RCs

[0284] To demonstrate the scalability and cost-efficiency of PSE and MPF for large cohort studies, a dataset of 2,561 run-control (RC) injections was processed, comprising two pooled human peptide mixtures analyzed across four Orbitrap Astral instruments over an eight-month interval. These stable peptide mixtures provide an ideal test bed for measuring computational scalability independent of biological variability (FIG. 27).

[0285] The dataset was successfully processed by both engines; however, DIA-NN’ s single-node execution model limits its ability to scale far beyond this study, particularly for high-depth sample sets where memory requirements become prohibitive. In contrast, workloads are distributed by PSE and MPF across many nodes with near-linear scaling, allowing thousands of acquisitions to be processed in parallel. Processing of all 2,561 files search was completed by PSE and MPF in 3.99 and 2.42 hours of real “wall” time, with end-to-end cost of approximately $111 and $110 for library-free and MBR analyses, respectively. Approximately 1.92x more wall time and at least 9.5x higher cloud costs were required by DIA-NN, running in the cloud-distributed Scalable MBR workflow previously described, to complete the same analyses. Wall time analysis reflects limitations of cloud resource availability, as well as start-up and orchestration overhead, resulting in lower fold improvement as compared to costs. Overall cost of execution reflects both the amount of provisioned compute resources and the time for which they were utilized, demonstrating the significant improvement in computational efficiency provided by PSE and MPF. For large datasets, costs are dominated by parallel processing of individual acquisitions, where fewer resources can be used by PSE while still executing more quickly.Example 7: Local and Cloud Execution Environments

[0286] PSE and MPF are designed for a wide range of environments, from a single computer to a cluster of many nodes. In this work, either a local execution or cloud execution environment was employed. The former runs all calculations in a single Docker container that contains PSE and MPF.WSGR Docket No. 53344-800.601The latter employs containerized PSE searches running on a Kubemetes cluster and an auto-scaling Spark cluster for MPF, each deployed within AWS. These environments differ significantly in capabilities and availability of compute resources, but there were only very minor differences in the code and configurations used to execute workflows. Other than the choice of search plugin used in the MPF workflow, the same MPF modules and configurations were used in both environments.

[0287] For local execution, a Docker image was prepared containing PSE and MPF, as well as runtime dependencies such as Python and Spark. For each analysis, a container was configured with appropriate volumes, and a TOML file passed to the MPF CLI. PSE was run sequentially on all input MS files using the local execution plugin described above. PSE and MPF output Parquet files were written to local storage. For local execution analyses presented in this manuscript, a c7a.8xlarge AWS instance with 32 CPUs and 64 GiB of memory was employed. The same instance type was used to perform DIA-NN analyses.

[0288] Cloud execution requires additional setup and orchestration, which was automated within an AWS environment. This system took a list of input MS files (stored in AWS S3) and a TOML configuration and set up an auto-scaling Spark cluster using Databricks. Necessary Python libraries, including MPF and the cloud PSE search plugin were installed in the cluster, and the MPF workflow executed using the provided parameters. PSE searches were run using the cloud search plugin, which executed them in parallel, using one container per MS file in an AWS EKS cluster. PSE results were written to S3 and read using Spark for subsequent workflow steps. MPF results were also written to S3.

[0289] PSE / MPF cloud execution analyses presented in this manuscript used the following configuration:

[0290] Spark Cluster:

[0291] 1 i3.2xlarge driver node (8 CPU / 61 GiB memory)

[0292] 1-20 i3.2xlarge worker nodes (8-160 CPU / 61-1220 GiB memory)

[0293] PSE searches (each MS file): 8 CPU / 32 GiB memory

[0294] The use of autoscaling ensured that only resources necessary for execution were allocated to the Spark cluster. The utilized configuration provides sufficient resources to scale to experiments of thousands of files with reasonable performance, without incurring unnecessary costs for smaller runs. While the number of nodes used for each analysis was not specifically tracked, aggregate costs for execution of each analysis were collected from AWS.

[0295] For comparison purposes, a cloud-distributed DIA-NN pipeline was also evaluated, running either in single-pass mode, or using an optimized MBR pipeline, termed “Scalable MBR”. While this pipeline also employs parallel processing of individual files where possible, it requires that results be collected onto a single processing node at one or more points during processing. Scalable MBR reduces the resource requirements but does not permit the same level of distributed processingWSGR Docket No. 53344-800.601or autoscaling possible with MPF. For DIA-NN cloud analyses presented here all steps were performed in containers running on AWS EKS with the following resource allocations:

[0296] DIA-NN searches (each MS file): 16 CPU / 64 GiB memory

[0297] Global processing (generation of MBR library, global FDR estimation and reporting): 74 CPU, 750 GiB memory

[0298] Quant normalization: 48 CPU, 512 GiB memory

[0299] Quant rollup: 64 CPU, 750 GiB memory

[0300] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

WSGR Docket No. 53344-800.601CLAIMS WHAT IS CLAIMED IS:

1. A method, comprising:receiving a plurality of retention times from a mass spectrometer;processing the plurality of retention times using a model configured to calibrate a retention time of the plurality of retention times, wherein the model configured to calibrate the retention time comprises a linear discriminant analysis (LDA);determining a first protein structure match based on the processed plurality of retention times; andtraining a machine learning model to output a second protein structure match, wherein the machine learning model is trained on a dataset comprising the first protein structure match and the plurality of processed retention times.

2. The method of claim 1, wherein the mass spectrometer is configured for data independent acquisition (DIA), and wherein receiving the plurality of retentiontimes comprises receiving (directly or indirectly) a data file generated by the mass spectrometer.

3. The method of claim 1, wherein the received plurality of retention times are stored in a database having a tree structure.

4. The method of claim 3, wherein the database has a R-tree structure (e.g., an original R-tree structure, a packed R-tree structure, or a similar database structure).

5. The method of claim 1, wherein the first protein structure match comprises an identified peptide and a corresponding false discovery rate (FDR).

6. The method of claim 1, wherein the second protein structure match comprises the first protein structure match and an optimized FDR.

7. The method of claim 1, wherein the model configured to calibrate the retention time comprises a retention time (RT) alignment model and an RT window filter.

8. The method of claim 7, wherein the RT window filter is derived from the RT alignment model.

9. The method of claim 1, wherein the model configured to calibrate the retention time is trained on a dataset comprising empirical retention data and predicted retention data, wherein the predicted retention data is obtained from a target library, and wherein the targets in the target library are peptides.

10. The method of claim 1, wherein receiving the plurality of retention times further comprises receiving a plurality of mass-to-charge (m / z) values, wherein each retention time of the plurality of retention times is paired with a corresponding m / z value of the plurality of m / z values.

11. The method of claim 10, further comprising:computing a mass calibration curve; andWSGR Docket No. 53344-800.601recalibrating the received plurality of corresponding m / z values using the computed mass calibration curve,wherein computing the mass calibration curve and recalibrating the received plurality of corresponding m / z values is performed after calibrating the plurality of retention times.

12. The method of any one of claims 1 to 11, wherein determining the first protein structure match is based on the processed plurality of retention times and the recalibrated plurality of corresponding m / z values.

13. The method of any one of claims 1 to 11, wherein receiving the plurality of retention times further comprises receiving a plurality of signal intensities, wherein each retention time of the plurality of retention times is paired with a corresponding signal intensity of the plurality of signal intensities.

14. The method of claim 13, wherein determining the first protein structure match is based on the processed plurality of retention times, the plurality of corresponding signal intensities, and, optionally, the recalibrated plurality of corresponding m / z values.

15. The method of any one of claims 1 to 11, wherein the machine learning model comprises a neural network, wherein the second protein structure match output by the machine learning model comprises an optimized FDR.

16. The method of any one of claims 1 to 11 further comprising:determining a plurality of first protein structure matches, wherein the machine learning model is trained on the plurality of first protein structure matches, the plurality of processed retention times, and the recalibrated plurality of corresponding m / z values.

17. The method of claim 16, wherein each protein structure match of the plurality of first protein structure matches has an FDR that is below a pre-selected threshold, wherein the preselected FDR threshold for the first protein structure matches is 50% FDR.

18. The method of claims 16, further comprising: outputting a list comprising a plurality of second protein structure matches, wherein each second protein structure match passes a preselected FDR threshold, and wherein the pre-selected FDR threshold for the plurality of second protein structure matches is an FDR of 1% or less.

19. A method for protein identification and / or quantification by mass spectrometry (MS) experiments, comprising:obtaining a plurality of mass spectra utilizing data-independent acquisition (DIA); processing by linear discriminant analysis (LDA) input data representing complex biological samples to build a model that correlates predicted retention times with empirical values;applying a retention time window filter based on the model generated by LDA; recalibrating mass-to-charge (m / z) values for all ions in the plurality of mass spectra using average mass calibration curves across batches; andWSGR Docket No. 53344-800.601determining peptide-spectrum matches (PSMs) based at least in part on the model generated by LDA, and wherein determining the PSMs comprises performing LDA.

20. The method of claim 19, wherein the method further comprises a high-throughput proteomics workflow for processing at least 2500 data files comprising the plurality of mass spectra within a 24-hour period, and wherein the method processes a population-scale proteomics study with an FDR of less than about 50%.

21. The method of claim 19, wherein a false discovery rate of the method is less than 5%.

22. The method of any one of claims 19 to 21, wherein the LDA processing is performed individually on each MS data file comprising mass spectra of the plurality of mass spectra.

23. The method of claim 22, wherein the LDA processing comprises a pairwise analysis of peptide target spectra and corresponding decoy spectra, wherein the peptide target spectra are a subset of peptide target spectra from a peptide target library.

24. The method of claim 23, wherein the peptide target spectra, the decoy spectra, and / or the peptide target library is / are stored in an accessible memory database having an R-tree structure, such as an original R-tree or a packed R-tree.

25. The method of claim 23, wherein the peptide target spectra and the corresponding decoy spectra are represented by multi-dimensional feature vectors, wherein the multi-dimensional feature vectors comprise at least 50 features.

26. The method of any one of claims 19 to 21, wherein the average mass calibration curves are generated using data output from the LDA processing.

27. The method of any one of claims 19 to 21, wherein determining the PSMs comprises generating feature vectors representative of each peptide target spectrum of a peptide target library, and scoring the feature vectors against mass spectra of the plurality of mass spectra having recalibrated m / z values, wherein each feature vector used for determining the PSMs comprises at least 100 features, and wherein determining the PSMs is further based at least in part on the retention time window filter and the recalibrated m / z values.

28. The method of any one of claims 19 to 21 further comprising normalization of ion counts or intensities prior to determining PSMs, and calculating precursor quantities for the determined PSMs.

29. The method of any one of claims 19 to 21 further comprising calculating precursor quantities for the determined PSMs, and wherein calculating a precursor quantity comprises summing detected fragment intensities over one or more scans.

30. A method for searching mass spectrometry data, the method comprising: obtaining a mass spectrometry (MS) data file comprising retention times and corresponding mass-to-charge (m / z) values;WSGR Docket No. 53344-800.601performing a first search of the MS data file using a first plurality of peptide target spectra from a library of peptide targets, thereby identifying a first set of peptide-spectrum matches (PSMs);calculating an optimized retention time (RT) window filter using predicted retention times for target peptides represented in the first set of PSMs and their corresponding observed retention times in the MS data file;computing a mass calibration curve using predicted m / z values for the target peptides in the first set of PSMs and corresponding observed m / z values in the MS data file;recalibrating the m / z values in the MS data file using the computed mass calibration curve, thereby generating a recalibrated MS data file;calculating optimized mass tolerances by performing a series of iterative searches on the recalibrated MS data file, using a second plurality of peptide target spectra from the peptide target library;scoring the peptide target library against the recalibrated MS data file using the optimized RT window filter and the optimized mass tolerances, thereby generating a set of optimized PSMs;calculating optimized false discovery rates (FDRs) for the optimized PSMs in the set of optimized PSMs using a machine learning algorithm.

31. The method of claim 30, further comprising outputting a listof optimized PSMs having optimized FDRs below a pre-selected FDR threshold (e.g., 1%), wherein the list of optimized PSMs comprises a feature vector and / or feature vector score, and related metadata for each optimized PSM in the list.

32. The method of claim 30, wherein the MS data file corresponds to a single MS injection, and wherein the MS data file has an mzML or Parquet file format.

33. The method of any one of claims 30 to 32, wherein the first plurality of peptide target spectra corresponds to about 0.1% to about 25% of the peptide targets in the peptide target library.

34. The method of any one of claims 30 to 32, wherein the first plurality of peptide target spectra comprises about 20,000 peptide target spectra.

35. The method of any one of claims 30 to 32 further comprising:obtaining the peptide target library; andstoring the peptide target library in a memory database having an R-tree structure (e.g., an original R-tree or a packed R-tree structure).

36. The method of claim 35, wherein each peptide target spectrum in the first plurality of peptide target spectra is represented as a vector having a plurality of features, wherein each vector representing a peptide target spectrum comprises at least 50 features.

37. The method of claim 36 further comprising:WSGR Docket No. 53344-800.601prior to building an RT alignment model, generating a decoy spectrum corresponding to each peptide target spectrum in the first plurality of peptide target spectra, wherein generating a decoy spectrum corresponding to a target spectrum comprises determining a spectrum corresponding to a mutated peptide sequence derived from a peptide sequence corresponding to the peptide target spectrum.

38. The method of claim 37, wherein decoy spectrum is represented as a vector having a plurality of features, wherein each decoy vector representing a peptide target spectrum comprises at least 50 features.

39. The method of claim 37, wherein building the RT alignment model comprises scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file, wherein scoring corresponding peptide target spectrum-decoy spectrum pairs against the MS data file comprises linear determinant analysis (LDA).

40. The method of claim 39, wherein scoring corresponding peptide target spectrumdecoy spectrum pairs against the MS data file comprises, for each potential PSM, generating a score for each of the peptide target spectrum and the corresponding decoy spectrum and determining, based upon a difference between those scores, whether the peptide target spectrum has a PSM with the MS data file.

41. The method of any one of claims 30 to 32, wherein calculating the optimized RT window comprises using splines generated from a subset of PSMs in the first set of PSMs, wherein each PSM in the subset of PSMs meets a pre-selected false discovery rate (FDR) of 5%, and wherein the subset of PSMs comprises at least 200 PSMs.

42. The method of any one of claims 30 to 32, wherein the mass calibration curve is computed a least squares regression to fit a polynomial, and wherein the polynomial is a second order polynomial .

43. The method of any one of claims 30 to 32, wherein the searches in the series of searches of the recalibrated MS data file use an ascending set of values for mass tolerance.

44. The method of any one of claims 30 to 32, wherein scoring the target library against the recalibrated MS data file comprises performing LDA, wherein scoring the target library against the recalibrated MS data file comprises generating feature vectors representative of each peptide target spectrum from the peptide target library, and scoring the feature vectors against the recalibrated MS data file, and wherein each feature vector comprises at least 100 features.

45. The method of any one of claims 30 to 32, wherein scoring the target library against the recalibrated MS data file further comprising calculating precursor quantities for the set of optimized PSMs, and wherein calculating a precursor quantity comprises summing detected fragment intensities over one or more scans.WSGR Docket No. 53344-800.60146. The method of any one of claims 30 to 32, wherein the machine learning algorithm is a neural network.

47. The method of claim 46, wherein the neural network comprises a classifier having three fully-connected layers and a configurable hidden layer width using ReLU activation, and wherein the neural network is trained using binary cross-entropy loss and multi-fold cross-validation.

48. The method of any one of claims 30 to 32 further comprising:obtaining a plurality of MS data files, each comprising retention times and corresponding m / z values; andsearching each of the MS data files in the plurality of MS data files,wherein the plurality of MS data files comprises at least 2500 MS data files.

49. The method of claim 48, wherein each MS data file in the plurality of MS data files is stored and / or searched in a single node (e.g., a single computer or a single processing unit in a distributed computing environment), and wherein each MS data file in the plurality of MS data files is searched independently and / or wherein the MS data files in the plurality of MS data files are searched in parallel.

50. The method of any one of claim 48, wherein each MS data file in the plurality of MS data files is searched independently and / or wherein the MS data files in the plurality of MS data files are searched in parallel.

51. A method for analyzing a biological sample comprising:processing a mass spectrometry data file using a library of predicted peptide spectra to identify i) a plurality of protein structure matches; ii) a first set of scores associes with the plurality of protein structure matches; and iii) a second set of scores associated with a plurality of decoy matches;dividing the first set of scores and the second set of scores into a plurality of blocks, each block containing a subset of the protein structure matches;computing, for each block, a count of the number of protein structure matches and a count of the number of decoy matches within that block; andusing the counts of the number of protein structure matches and the number of decoy matches within each block to estimate a false discovery rate (FDR) for each block; and aggregating the FDR estimates across the plurality of blocks to yield an overall FDR estimate for the plurality of protein structure matches,wherein the computing is performed using a distributed computing environment comprising a plurality of nodes.WSGR Docket No. 53344-800.60152. The method of claim 51, wherein the method comprises calculating, for each block, an upper-bound q-value estimate and a lower-bound q- value estimate based on the first set of scores and the second set of scores within that block.

53. The method of claim 51 or 52, wherein the number of blocks is between 20,000 and 100,000, or between 30,000 and 80,000.

54. The method of claim 51 or 52, wherein the method further comprises selecting an acceptable relative error for quantization, and wherein the acceptable relative error for quantization is less than 0.0005.

55. A method, comprising:receiving, at a plurality of processing units, mass spectrometry data from at least one mass spectrometer;processing the mass spectrometry data, wherein the processing comprises calibrating a retention time of the mass-spectrometry data; anddetermining a plurality of protein structure matches.

56. The method of claim 55, wherein the processing comprises performing the method of any one of claims 1 to 11, 30 to 32, 51 or 52.

57. The method of claim 34 or 35, wherein: the plurality of protein structure matches has a false match rate of less than about 0.5%; wherein a false match rate at a precursor level is less than about 1%; and / or wherein a false match rate at a protein group level is less than about 1%.

58. A non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to perform operations comprising a method of any one of claims 1 to 11, 30 to 32, 51, or 52.

59. A system for performing quantitative DIA proteomics, comprising:a computer configured to perform the methods of any one of claims 1 to 11, 30 to 32, 51, or 52;a data storage system for managing and processing one or more mzML files, each independently comprising data-independent analysis mass spectra of one or more biological samples; anda distributed computing environment (e.g., Apache Spark) for scalable execution of the data analysis pipeline.

60. The system of claim 59, wherein the distributed computing environment is scalable to analyze at least 1000 samples per day.