Methods and systems for spectrometric analysis
The technologies address the challenges of analyzing complex deconvoluted spectra by employing systematic methods for data processing and visualization, thereby enhancing the accuracy and confidence of spectral analysis in high-throughput mass spectrometry.
Patent Information
- Application Number
- PCT/US2024/053632
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-12
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
Current methods for spectrometric analysis, particularly in high-throughput mass spectrometry, face challenges in efficiently processing and analyzing large datasets of complex deconvoluted spectra, often relying on estimates due to the lack of orthogonal supporting data, which leads to poorly constrained search spaces and increased false annotations.
The described technologies involve systems and methods for analyzing and displaying mass spectra, including obtaining digital data from biological samples, performing deconvolution and mass subtraction operations, generating a global mass axis, and producing intermediate data for clustering, factorization, or sequence-spectra predictions to generate graphical representations of the output data.
These methods improve the accuracy and confidence of spectral analysis by providing a systematic approach to data processing, reducing false annotations, and enabling the visualization of complex data sets, thus facilitating the identification of specific qualities or properties in biological samples like monoclonal antibodies.
Smart Images

Figure US2024053632_08052025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR SPECTROMETRIC ANALYSISFIELD
[0001] The present disclosure relates to technologies, including methods and systems, for spectrometric analysis.BACKGROUND
[0002] Mass spectrometry can be used for the analysis of the mass of biological samples and is well suited for in-depth analysis or screening of proteins, such as monoclonal antibodies (mAbs). Broad screening of proteins by intact mass analysis (with or without enzymatic or chemical modification) using techniques that measure mass distributions (e.g., mass spectrometry) produces a full underlying mass distribution of a sample of proteins from which heterogeneities in the sample can be identified. These techniques can likewise be used for protein mixtures that are chromatographically resolved. Examples of heterogeneities in proteins, such as mAbs, include (but are not limited to): glycoform distribution, N-terminal modification, C-terminal lysine retention, and leader sequence retention. These data may aid in identifying and selecting, from a set of samples, mAbs having one or more specific qualities or properties. The ability to scale such methods for high-throughput mass spectrometry (HT-MS) is limited by the time and resources needed to critically analyze hundreds or more of individual spectra. The rate at which data can be generated in such a HT-MS system vastly exceeds the rate at which that data can be processed and utilized.SUMMARY
[0003] The analysis of complex deconvoluted spectra is challenging, particularly where supporting experimental data, such as peptide mapping and released glycan analysis, is not available. Without such orthogonal supporting data, analysis may rely on estimates derived from frequently observed post-translational modifications from expression systems, protein modalities, expression conditions, etc., that are expected to show similarity with those of the protein sample under analysis. This may result in a search space, used to describe the mass distribution / deconvoluted spectrum, that is poorly constrained. For instance, where the search space is defined as the allowed post-translational modifications that can be combinatorically applied to adjust the mass of the expected sequence, the number of falseannotations may increase as search space expands, while the ability to annotate spectra with high confidence decreases.
[0004] The technologies described in this specification include systems and methods for analyzing and displaying mass spectra. A system may include a computing device comprising a processor, a memory and a set of instructions that, when executed by the processor, cause the processor to perform one or more operations.
[0005] In an aspect, the methods and system operations include obtaining digital data comprising a set of mass spectra from a set of biological samples and performing a deconvolution operation on each member of the set of mass spectra. The methods and system operations include performing a mass subtraction operation on the deconvoluted mass spectra to obtain a set of adjusted mass distributions, comprising subtraction of known or expected mass from the measured mass, producing adjusted mass values. The methods and system operations include generating a global mass axis and generating a rectangular matrix of the set of adjusted mass distributions wherein the global mass axis is shared by the adjusted mass distributions. The methods and system operations include generating, from the rectangular matrix, intermediate data, the intermediate data comprising at least one of correlative metrics, similarity metrics, or statistical data. The methods and system operations include performing at least one operation of clustering, factorization, unsupervised clustering, or sequencespectra predictions to obtain a set of output data. The methods and system operations include generating and causing to display a graphical representation of the output data.
[0006] In another aspect, the methods and system operations include obtaining digital data comprising a set of mass spectra from a set of biological samples and performing a deconvolution operation on each member of the set of mass spectra. The methods and system operations include performing an alignment operation on the deconvoluted mass spectra to obtain a set of adjusted mass distributions. The methods and system operations include generating a global mass axis and generating a rectangular matrix of the set of adjusted mass distributions wherein the global mass axis is shared by the adjusted mass distributions. The methods and system operations include generating, from the rectangular matrix, intermediate data, the intermediate data comprising at least one of correlative metrics, similarity metrics, or statistical data. The methods and system operations include performing at least one operation of clustering, factorization, unsupervised clustering, or sequence-spectra predictions to obtain a set of output data. The methods and system operations include generating and causing to display a graphical representation of the output data.
[0007] In another aspect, the methods and system operations include obtaining digital data comprising a set of mass spectra from a set of biological samples and performing a deconvolution operation on each member of the set of mass spectra. The methods and system operations include performing a mass subtraction operation on the deconvoluted mass spectra to obtain a first set of adjusted mass distributions, comprising subtraction of known or expected mass from the measured mass, producing adjusted mass values. The methods and system operations include generating a global mass axis and generating a first rectangular matrix of the set of adjusted mass distributions wherein the global mass axis is shared by the adjusted mass distributions. The methods and system operations include identifying, from the set of adjusted mass distributions, a highly correlated sample group of mass distributions and performing an alignment operation of non-correlated samples against the mass-adjusted, correlated sample group to obtain a second set of adjusted mass distributions. The methods and system operations include outputting a mass shift to identify unknown sequence mass and generating a second rectangular matrix of the second set of adjusted mass distributions. The methods and system operations include generating, from the second rectangular matrix, intermediate data, the intermediate data comprising at least one of correlative metrics, similarity metrics, or statistical data. The methods and system operations include performing at least one operation of clustering, factorization, unsupervised clustering, or sequencespectra predictions to obtain a set of output data. The methods and system operations include generating and causing to display a graphical representation of the output data.
[0008] In some embodiments, the alignment operation comprises performing a globally optimized alignment operation wherein a set of mass distributions are aligned such that the differences of a mass distribution against all other mass distributions in the set is minimized.
[0009] In some embodiments, the alignment operation comprises aligning all mass distributions to a single mass distribution. In some embodiments, the alignment operation comprises aligning all mass distributions against all mass distributions using a global optimization approach. In some embodiments, the alignment operation includes performing an operation comprising comparing a calculated mass shift value to an expected mass shift value.
[0010] In some embodiments, generating the global mass axis comprises finding all unique mass values across the set of mass spectra, sorting the mass values in increasing order, and using said order of mass values as the global mass axis. In some embodiments, generatingthe global mass axis comprises selecting the highest numerical accuracy mass axis from the set of mass spectra.
[0011] In some embodiments, generating the intermediate data comprises generating and displaying a heat map or a dendrogram. In some embodiments, generating the intermediate data comprises performing a network analysis operation. In some embodiments, displaying a graphical representation of the output data comprises displaying one of a cluster visualization, a dimensionality reduced visualization, or an underlying factor visualization.
[0012] In some embodiments, the biological samples comprise material from biobetters or biosimilars or a reference sample, or a combination thereof. In some embodiments, the biological samples comprise material from a set of cell lines or material from different biomolecule purification or manufacturing processes.
[0013] Another aspect of the present disclosure relates to a computer-implemented method comprising: obtaining mass distributions corresponding to samples produced using a high- throughput expression system; from the mass distributions, determining aligned mass distributions represented on a global mass axis; and causing display of a graphical representation corresponding to the aligned mass distributions.
[0014] In some embodiments, determining the aligned mass distributions comprises subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions.
[0015] In some embodiments, determining the aligned mass distributions comprises aligning the mass distributions to a single mass distribution of the mass distributions.
[0016] In some embodiments, determining the aligned mass distributions comprises minimizing the differences of a mass distribution against all others of the mass distributions.
[0017] In some embodiments, determining the aligned mass distributions comprises: determining adjusted mass distributions by subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions; after the subtracting, generating the global mass axis; generating a first matrix including the adjusted mass distributions represented on the global mass axis; identifying, from the first matrix, a correlated group of the adjusted mass distributions; performing an alignment operation of a non-correlated group of adjusted mass distributions against the correlated group; and generating a second matrix including the non-correlated group aligned with the correlated group.
[0018] In some embodiments, determining the aligned mass distributions comprises: obtaining adjusted mass values by subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions; and performing a modification adjustment on the adjusted mass values.
[0019] In some embodiments, the modification adjustment corresponds to a chemical modification of a protein, a polynucleotide, a virus, a virus-like particle, a nanoparticle, a peptide, a lipid, a saccharide, or a metabolite.
[0020] In some embodiments, the modification adjustment corresponds to one or more of glycosylation, de-glycosylation, phosphorylation, acetylation, methylation, ubiquitination, sumoylation, lipidation, proteolysis, deamidation, hydroxylation, protein labeling, base insertion, base deletion, or base substitution.
[0021] In some embodiments, the computer-implemented method further comprises generating the global mass axis by finding all unique mass values across the mass distributions, sorting the mass values in increasing order, and using the increase order as the global mass axis.
[0022] In some embodiments, the computer-implemented method further comprises generating the global mass axis by selecting the highest numerical accuracy mass axis from the mass distributions.
[0023] In some embodiments, causing display of the graphical representation comprises causing display of one or more of a cluster visualization, a dimensionality reduced visualization, or an underlying factor visualization.
[0024] In some embodiments, a first portion of the mass distributions is obtained at a first time point of a sample process and a second portion of the mass distributions is obtained at a second time point of the sample process.
[0025] In some embodiments, the first portion and the second portion are obtained from samples of the same substance.
[0026] In some embodiments, the samples are biological samples.
[0027] Another aspect of the present disclosure relates to a system comprising: at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, further cause the system to: obtain mass distributions corresponding to samples produced using a high-throughput expression system; from the mass distributions, determine aligned mass distributions represented on a global mass axis; and cause display of a graphical representation corresponding to the aligned mass distributions.
[0028] In some embodiments, the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to subtract the known or expected masses of the samples from the measured masses of the samples in the mass distributions.
[0029] In some embodiments, the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to align the mass distributions to a single mass distribution of the mass distributions.
[0030] In some embodiments, the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to minimize the differences of a mass distribution against all others of the mass distributions.
[0031] In some embodiments, the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to: determine adjusted mass distributions by subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions; after the subtracting, generate the global mass axis; generate a first matrix including the adjusted mass distributions represented on the global mass axis; identify, from the first matrix, a correlated group of the adjusted mass distributions; perform an alignment operation of a non-correlated group of adjusted mass distributions against the correlated group; and generate a second matrix including the non-correlated group aligned with the correlated group.
[0032] In some embodiments, the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to: obtain adjusted mass values by subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions; and perform a modification adjustment on the adjusted mass values.
[0033] In some embodiments, the modification adjustment corresponds to a chemical modification of a protein, a polynucleotide, a virus, a virus-like particle, a nanoparticle, a peptide, a lipid, a saccharide, or a metabolite.
[0034] In some embodiments, the modification adjustment corresponds to one or more of glycosylation, de-glycosylation, phosphorylation, acetylation, methylation, ubiquitination,sumoylation, lipidation, proteolysis, deamidation, hydroxylation, protein labeling, base insertion, base deletion, or base substitution.
[0035] In some embodiments, the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to generate the global mass axis by finding all unique mass values across the mass distributions, sorting the mass values in increasing order, and using the increase order as the global mass axis.
[0036] In some embodiments, the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to generate the global mass axis by selecting the highest numerical accuracy mass axis from the mass distributions.
[0037] In some embodiments, the instructions for causing display of the graphical representation further comprise instructions that, when executed by the at least one processor, further cause the system to cause display of one or more of a cluster visualization, a dimensionality reduced visualization, or an underlying factor visualization.
[0038] In some embodiments, a first portion of the mass distributions is obtained at a first time point of a sample process and a second portion of the mass distributions is obtained at a second time point of the sample process.
[0039] In some embodiments, the first portion and the second portion are obtained from samples of the same substance.
[0040] In some embodiments, the samples are biological samples.INCORPORATION BY REFERENCE
[0041] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Novel features of the technologies described in this specification are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present technologies will be obtained by reference to the following detailed descriptionthat sets forth illustrative embodiments, in which the principles of the technologies are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0043] FIG. 1A is a flow chart illustrating an example workflow using mass adjustment to subtract known contribution of molecular compositions (e.g., protein sequences) on mass spectra and to generate a global mass axis;
[0044] FIG. IB is a flow chart illustrating an alternative example workflow using mass adjustment to subtract known contribution of molecular compositions (e.g., protein sequences) on mass spectra and to generate aligned mass spectra / distributions;
[0045] FIG. 2A is a flow chart illustrating an example workflow using global alignment to align spectra and to generate a global mass axis;
[0046] FIG. 2B is a flow chart illustrating an alternative example workflow using global alignment to generate aligned mass spectra / distributions;
[0047] FIG. 3A is a flow chart illustrating an example workflow using an iterative approach where mass adjustments are made based on known differences (i.e., the expected protein sequence) followed by pairwise calculation of mass distribution similarity using the adjusted mass distributions to align spectra and to generate a global mass axis;
[0048] FIG. 3B is a flow chart illustrating an alternative example workflow using an iterative approach to generate aligned mass spectra / distributions, where mass adjustments are made based on known differences (i.e., the expected sample composition, e.g., protein sequence) followed by pairwise calculation of mass distribution similarity;
[0049] FIG. 4 is a diagrammatic representation of an example system according to an example implementation of the technologies described in this specification;
[0050] FIG. 5A is a graph showing example mass spectra obtained by deconvolution of the received spectrometry data from two samples; FIG. 5B is a data table showing the value matrix of the data of the graph in FIG. 5A;
[0051] FIG. 6A is a graph showing the example mass spectra from FIG. 5A after mass adjustment; FIG. 6B is a data table showing the value matrix of the data of the graph in FIG.6A;
[0052] FIG. 7A is a graph showing the example adjusted mass spectra from FIG. 6A aligned to a global mass axis; FIG. 7B is a data table showing the value matrix of the data of the graph in FIG. 7A;
[0053] FIGS. 8A and 8B are graphs showing example mass spectra obtained from deconvolution of received spectrometry data from a plurality of samples;
[0054] FIG. 9A is a graph of an example set of mass spectra that have been mass adjusted so that the position of peaks and troughs along the adjusted mass axis represent mass deviation from the expected sequence mass; FIG. 9B is a graph of an example set of mass spectra that have been aligned and where the signal intensity (y-axis) values have been normalized;
[0055] FIG. 10 is a graph showing an example adjusted mass distribution of Immunoglobulin G (IgG) antibodies obtained by deconvolution of the raw mass spectrometry data;
[0056] FIG. 11 is a flow chart illustrating an example workflow using mass adjustment to subtract known contribution of molecular compositions (e.g., protein sequences) on mass spectra and using modification adjustment to generate a global mass axis;
[0057] FIG. 11B is a flow chart illustrating an alternative example workflow using mass adjustment to subtract known contribution of molecular compositions (e.g., protein sequences) on mass spectra and using modification adjustment to generate aligned mass spectra / distributions
[0058] FIG. 12A is a graph showing the whole mass range of the spectra shown in FIG. 10 after modification adjustment; FIG. 12B is a graph showing the range indicated by the black bars in FIG. 12A;
[0059] FIG. 13 is a flow chart illustrating an example analysis workflow from mass-aligned deconvoluted spectra to calculation of pairwise statistics to conversion to a pairwise square matrix;
[0060] FIGS. 14A and 14B show graphs and flow charts illustrating example visualization workflows from square matrices to seriation to clustering;
[0061] FIG. 15A is a graph of an example pairwise matrix of a comparison of an example protein data set; FIG. 15B is a graph illustrating DBSCAN clustering of the same data;
[0062] FIG. 16 is a graph illustrating 150 deconvoluted mass spectra of six example groups of simulated proteins;
[0063] FIG. 17A is a graph illustrating the 150 spectra from FIG. 16 after mass adjustment; FIG. 17B are three graphs of the deconvoluted spectra of groups 1, 2, and 3.
[0064] FIG. 18A is a graph of an example pairwise matrix of a comparison of the 150 spectra from FIG. 16 with group labels. FIG. 18B is a graph illustrating the clustering according to the actual group membership (ground truth data);
[0065] FIG. 19A is illustrating DBSCAN clustering of the 150 spectra from FIG. 16 with an epsilon value of 0.01; FIG. 19B is illustrating DBSCAN clustering of the 150 spectra from FIG. 16 with an epsilon value of 0.005;
[0066] FIG. 20 is a schematic diagram illustrating an example workflow for training a machine-learning model for spectrometry data grouping using the alignment and clustering technologies described in this specification;
[0067] FIG. 21 is a flow chart illustrating an example workflow for training a machinelearning model for spectrometry data grouping using the alignment and clustering technologies described in this specification;
[0068] FIG. 22 is a schematic diagram illustrating an example workflow for spectrometry data grouping by executing a machine-learning model trained using the alignment and clustering technologies described in this specification; and
[0069] FIG. 23 is a flow chart illustrating an example workflow for spectrometry data grouping by executing a machine-learning model trained using the alignment and clustering technologies described in this specification.DEFINITIONS
[0070] The term Mass spectrum (mass spectra) as used in this specification can include the following definition: the mass spectrum is the representation of a data output of a mass spectrometer and includes ensemble collection of data of ions whose mass-to-charge (m / z) ratios had been determined. The data are displayed as intensity (or count) as a function of m / z and are referred to as the mass spectrum.
[0071] The term Mass distribution as used in this specification can include the following definition: the mass distribution is a plot of intensity / count / abundance (e.g., paired data comprising signal intensity) as a function of molecular mass. In mass spectrometry, this is the deconvoluted spectrum. In charge detection mass spectrometry, this spectrum can be referred to as the “True Mass Spectrum”. Additionally or alternatively to mass spectrometry, mass photometry can be used to obtain a mass distribution plot. Mass photometry is a bioanalytical technology that can measure molecular mass by quantifying light scattering from individual biomolecules in solution. Mass photometry can display contrast as a function of count.
[0072] The term Mass distribution alignment as used in this specification can include the following definition: mass distribution alignment refers to a process in which multiple mass distributions from one or more unique samples and / or one or more instruments (e.g., spectra, e.g., from multiple samples and multiple instruments) are aligned using a variety of methods described in this specification, e.g., as described in detail below.
[0073] The term Adjusted mass distribution as used in this specification can include the following definition: the resulting mass distribution after mass distribution alignment.
[0074] The term Global mass axis as used in this specification can include an axis (e.g., the x-axis) of an adjusted mass distribution and can be calculated such that corresponding intensity values along the global mass axis contain no missing values for any sample in the set.
[0075] The term Sample as used in this specification can include any physical matter whose mass distribution can be measured from a diverse set of measurement techniques. In some embodiments, a sample may include one or more biological materials. In some embodiments, a sample may include one or more non-biological materials. In some embodiments, a sample may include one or more biological macromolecules. In some embodiments, a sample may include one or more non-biological macromolecules. For example, a sample may include one or more deoxyribonucleic acid (DNA) molecules, one or more ribonucleic acid (RNA) molecules, one or more proteins, one or more antibodies, one or more lipids, and other like physical matter. While the present disclosure provides examples regarding the processing of samples containing proteins, one skilled in the art will appreciates that the same results would be expected if the samples contained other physical matter.DESCRIPTION
[0076] Analytical data processing, such as analysis of high-throughput mass spectrometry (HT-MS) data, requires scalable methods and data reduction strategies for data comparison and review that do not rely on individual data output files to be reviewed by humans. Described in this specification are technologies for spectrometric analysis, e.g., mass analysis, e.g., intact mass analysis on sample variants (e.g., produced using a plate-based high-throughput expression system). The technologies described in this specification are applicable to a variety of spectrometric analysis techniques and applications in astronomy, chemistry (e.g., drugs, pharmaceuticals, biologies), materials science (e.g., polymers, composites, metals), physics (e.g., particle physics), and biochemistry, e.g., the analysis ofproteins, polynucleotides (e.g., base insertion, base deletion, or base substitution), viruses, virus-like particles, nanoparticles, peptides, lipids, saccharides, metabolites (e.g. amino acids, small alcohols, amines, polycarboxylic acids, etc.), and / or other (e.g., biochemical) molecules. The technologies described in this specification can be used to analyze spectra in sets of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000 spectra.
[0077] In an example implementation of the technology, raw mass spectrometry data are deconvoluted to provide mass distributions. These mass distributions can be adjusted by removing the mass contribution from the anticipated molecule (e.g., protein sequence). In other words, it is assumed that the intended molecule (e.g., protein sequence) is present in each sample and, by removing the sequence mass, the spectra can be well aligned. This process produces a mass distribution representing the post-translational heterogeneity present in each sample, which can be compared across multiple molecules (e.g., proteins) - this process is notably applicable to protein variant sets whereby sequence variants are generated and expressed in the same expression system and where the differences in the heterogeneity profile are expected to be limited.
[0078] The technologies described in this specification differ from other processes that, e.g., are used to detect repetitious mass differences between components within a single sample or single spectrum. Some other processes require determination of mass differences between peaks within a spectrum (e.g., through application of sliding error windows to find positive and negative differences between peaks and generation of a mass difference grid) and do not include mass adjustment based on prior knowledge or assumption of mass. The technologies described in this specification, however, can provide analysis and classification of spectra across samples (rather than just determining repetitious mass differences between components within a single sample).Overview
[0079] FIG. 1A is a flow chart illustrating an example workflow using mass adjustment to subtract known contribution of molecular compositions (e.g., protein sequences) on massdistributions and to generate a global mass axis. Raw data is acquired (101) and deconvoluted to obtain the spectra (102), as described in detail herein below under the heading “Data Acquisition and Deconvolution. The expected mass of a molecule, e.g., a protein or biomolecule or macromolecular complex, etc., is hypothesized to be known a priori. The expected mass (Mcaic) is subtracted from the measured mass (Mobs) producing the adjusted mass (Madj) where Madj = Mcaic - Mobs (103). Thus, after data acquisition and deconvolution, the known or estimated mass of each analyte is subtracted from the corresponding spectrum, resulting in an alignment of the spectra as illustrated below. In some implementations, the technologies described in this specification utilize “subtraction of known differences” for mass adjustment. A global mass axis, where each x-value in the whole data set across spectra has a corresponding y-value, can be generated (104). Adjusted mass distributions can be normalized so their intensity values span a range from 0 to 1 and can then be compared pairwise using a variety of methods, e.g., Pearson’s r value, linear regression (slope, y-intercept), and / or the spectral contrast angle. From the globally aligned data a rectangular matrix is generated (105) that can be processed further as described below, e.g., using similarity metrics, correlative statistics (106), and / or visualization methods (107). In some example implementations, unsupervised clustering can be applied to the pairwise statistical metrics. Consensus spectra can be produced from these groupings. Visualization of data, e.g., clustered data and / or consensus data, can be in the form of ordered and / or unordered heat maps, dendrograms, and / or can be obtained using spectral averaging. Ultimately, group membership can be reported to determine pass / fail criteria to replace or supplement traditional verification methods based on annotation. Additional or alternative approaches to this general concept can include sliding window correlation approaches whereby individual samples are compared as described above or where a single consensus spectrum is used for comparison. Additional or alternative analysis methods and workflows are further described below.
[0080] Additionally or alternatively to mass adjustment as described above, mass distributions can be processed using unknown global alignment where a set of mass distributions are aligned such that the differences of a mass distribution against all other mass distributions in the set is minimized. This technique (“unknown global alignment” or “correlative alignment”) can be used, e.g., if mass of each analyte is not known and / or can be used for examining the similarity of heterogeneity profiles without regard of the underlying molecular composition (e.g., protein sequence).
[0081] FIG. IB is a flow chart illustrating an alternative example workflow using mass adjustment to subtract known contribution of molecular compositions (e.g., protein sequences) on mass distributions and to generate aligned mass distributions. Raw data is acquired (101) and deconvoluted (102) (e.g., depending on the data collection, to obtain mass spectra having corresponding mass distributions, as described in detail herein below under the heading “Data Acquisition and Deconvolution.”
[0082] Following data acquisition (101) and deconvolution (102), the mass spectra and corresponding mass distributions may undergo globally optimized alignment processes (108). As part of this, the mass spectra / distributions of the various samples may be aligned.
[0083] Adjusted mass spectra / distributions may be generated (109) from the (raw) mass spectra / distributions. The expected mass of a molecule, e.g., a protein or biomolecule or macromolecular complex, etc., may be hypothesized to be known a priori. The expected mass of a sample (Mcaic) may be subtracted from the measured mass (Mobs) of the sample, thereby producing the adjusted mass (Madj) for the sample, where Madj = Mcaic - Mobs. Thus, after data acquisition and deconvolution, the known or estimated mass of each analyte / sample may be subtracted from the corresponding measured mass of its mass spectrum / distribution, resulting in an adjusted mass spectrum / distribution for the sample. A consequence of this process is that the adjusted mass spectra / distributions for the different samples may be more aligned than the corresponding raw mass spectra / distributions (compare FIG. 5 A with FIG. 6A).
[0084] Further disclosure regarding generation of the adjusted mass spectra / distributions can be found herein below under the heading “Mass Distribution Alignment.”
[0085] After generating the adjusted mass spectra / distributions, a global mass axis, where each mass value (e.g., x-axis value) in the whole data set across mass spectra / distributions has a corresponding mass intensity value (e.g., y-axis value), can be generated (104). In other words, the global mass axis is obtained by causing each mass value (e.g., x-axis value) to have a corresponding mass intensity value (e.g., y-axis value) from each adjusted mass spectrum / distribution. Details of how the global mass axis may be generated are described below under the headings “Mass Distribution Adjustment” and “Generation of the Matrix.”
[0086] Aligned mass spectra / distributions may be generated (110) using the global mass axis and the adjusted mass spectra / distributions. Details of how the aligned mass spectra / distributions may be generated are described below under the heading “Mass Distribution Adjustment.”
[0087] The aligned mass spectra / distributions can optionally be normalized (111) to produce normalized mass spectra / distributions. In the normalized mass spectra / distributions, mass intensity values are normalized to span a range (e.g., from 0 to 1). Details of how this may be done are described below under the heading “Mass Distribution Adjustment.”
[0088] The aligned mass distributions (optionally normalized mass distributions) can thereafter undergo post-alignment processing (112). For example, the aligned mass distributions (optionally normalized mass distributions) can be compared pairwise and / or groupwise, using a variety of methods, e.g., Pearson’s r value, linear regression (slope, y- intercept), and / or spectral contrast angle. As another example, from the (globally) aligned mass distributions (optionally normalized mass distributions) a (e.g., rectangular) matrix may be generated, as described herein under the heading “Generation of the Matrix.” The matrix can be processed further, e.g., using similarity metrics, correlative statistics, and / or visualization methods. In some example implementations, unsupervised clustering can be applied to pairwise statistical metrics. Consensus spectra can be produced from these groupings.
[0089] Visualization of data, e.g., clustered data and / or consensus data, can be in the form of ordered and / or unordered heat maps (i.e., visualization tools that use colors to represent dataset value magnitudes), dendrograms (i.e., branching diagrams that show relationships between grouped entities based on their similarities), and / or can be obtained using spectral averaging (i.e., a method for calculating the average response of a system over a range of frequencies).
[0090] Ultimately, group membership can be reported to determine pass / fail criteria to replace or supplement traditional verification methods based on annotation. Additional or alternative approaches to this general concept can include sliding window correlation approaches whereby individual samples are compared or where a single consensus spectrum is used for comparison. Additional or alternative analysis methods and workflows are further described below.
[0091] Additionally or alternatively to mass adjustment as described with respect to FIGS. 1A and IB, aligned mass distributions can be produced using unknown global alignment where a set of mass distributions are aligned such that the differences of a mass distribution against all other mass distributions in the set is minimized. This technique (“unknown global alignment” or “correlative alignment”) can be used, e.g., if mass of each analyte is not knownand / or can be used for examining the similarity of heterogeneity profiles without regard of the underlying molecular composition (e.g., protein sequence).
[0092] FIG. 2 is a flow chart illustrating an example workflow using global alignment to align spectra and generate a global mass axis. Raw data is acquired (101) and deconvoluted to obtain the spectra (102). After data acquisition and deconvolution, a set of mass distributions are aligned such that the differences of a mass distribution (Mdist) against all other mass distributions in the set is minimized. This can be calculated by, e.g., (i) aligning all mass distributions to a single mass distribution, or (ii) aligning all mass distributions against all mass distributions using a global optimization approach (201). In some implementations, a mass shift (difference of a mass distribution) is compared to a known or expected value (201 A). An iterative optimization approach departing from an initial guess of expected mass and / or local maxima can be used employing one or more of, for example, genetic algorithms, simulated annealing, or particle swarm optimization techniques. A global mass axis, where each x-value in the whole data set across spectra has a corresponding y- value, can be generated (104). From the aligned data a rectangular matrix is generated (105) that can be processed further as described below, e.g., using similarity metrics, correlative statistics (106), and / or visualization methods (107).
[0093] FIG. 2B is a flow chart illustrating an alternative example workflow using global alignment to align spectra and generate a global mass axis. Raw data is acquired (101) and deconvoluted (102) to obtain mass spectra / distributions. After data acquisition and deconvolution, the mass spectra / distributions may undergo global optimized alignment processing (201). In the context of FIG. 2B, the mass spectra / distributions may be aligned such that the differences of a mass distribution (Mdist) against all other mass distributions in the set are minimized. This can be calculated by, e.g., (i) aligning (202a) all mass distributions to a single mass distribution, or (ii) aligning (202b) all mass distributions against all mass distributions using a global optimization approach. To align all mass distributions to a single mass distribution, for each mass distribution, the mass (e.g., x) axis may be slid (or “jittered”) until a similarity / correlation / root mean square deviation / etc. is maximized / minimized when comparing the selected mass distribution against the mass distribution under consideration. To align all mass distributions against all mass distributions, a set of values (one for each spectrum / sample), which represent the amount that each spectrum in the set should be adjusted, can be obtained. Then, an optimum set of values may be determined by minimizing / maximizing the similarity / correlation / root mean squaredeviation / etc. of all spectra against all spectra. This is most efficiently calculated by updating the entire set of adjustment values simultaneously.
[0094] In some implementations, a mass shift (difference of a mass distribution) is compared to a known or expected value (203). Two mass shifts exist. One is obtained from approach(es) (i) and / or (ii) in the above paragraph. The value of the mass shift can be compared to what the mass difference between samples could be. This lets the algorithm do the alignment of the spectra rather than depend on what is known about the sample. The result of the alignment algorithm can be compared to what is known about the sample.
[0095] If alignment is determined based on what known about the samples, spectra that do not align can be determined. In the context of steps 202a, 202b and 203, the alignment algorithm aligns everything and then looks for mass shifts that do not make sense. Finding those that do not make sense can be done by comparison to known values or by looking for outliers.
[0096] The iterative optimization approach of FIG. 2B, which departs from an initial guess of expected mass and / or local maxima, can employing one or more of, for example, genetic algorithms (i.e., methods of solving constrained and unconstrained optimization problems and which are based on natural selection), simulated annealing (i.e., a probabilistic technique for approximating the global optimum of a given function), or particle swarm optimization (i.e., a computational method that optimizes a problem by iteratively trying to improve a candidate solution with regard to a given measure of quality) techniques.
[0097] A global mass axis, where each mass value (e.g., x-axis value) in the whole data set across mass spectra has a corresponding mass intensity value (e.g., y-axis value), can be generated (104). Thereafter, the workflow of FIG. 2 A may include steps 110-112 of FIG. IB
[0098] Additionally or alternatively to mass or global alignment as described above, a combined, iterative anchored alignment process can be used. Some fraction of the samples may contain unintended / unexpected molecule (e.g., protein) whose composition (e.g., sequence) differs from the intended composition (e.g., sequence). For example, in the context when the samples contain proteins, the sequences may differ due to, e.g., truncations, fragmentations, substitutions, and / or DNA errors. Other workflows would require searching modifications for these possibilities of which there may be many; thus, the search space grows unacceptably large. The technologies described here obviate the need for searching said modifications.
[0099] FIG.3A is a flow chart illustrating an iterative approach where after data acquisition (101) and deconvolution (102), mass adjustments are made based on known differences (e.g., due to the expected protein sequence) (103) followed by pairwise calculation of mass distribution similarity using the adjusted mass distributions. A global mass axis, where each x-value in the whole data set across spectra has a corresponding y-value, can be generated (104). From the aligned data a first rectangular matrix is generated (105) that can be processed further as described below, e.g., using similarity metrics, correlative statistics (106). Some fraction of the sample will be well correlated - these represent the samples that (a) have similar heterogeneity profiles, and (b) where the known difference (e.g., the expected protein sequence) is present in the sample. These samples are identified (301) and / or removed from the data set and are then used as anchors against which all other samples are aligned by minimizing the residual difference between each poorly correlated sample and the set of well correlated samples (anchors) (302). The mass difference that minimizes the residual difference between the anchor samples and the sample under investigation is the hypothesized mass of the underlying molecule (e.g., protein) (303). This hypothesis is particularly strong if - after alignment - the similarity score (or correlation coefficient) is strong between the mass distribution under consideration and the anchor set. Similar to the other approaches, this method is well suited for protein variant sets with sequences containing a limited number of amino acid substitutions on a relatively constant backbone and expressed in the same expression system. The approach described herein is also well suited for the alignment of dissimilar data sets where dissimilar molecules (e.g., proteins) may be present and / or the approach can be used for untargeted clustering and aggregation of all data collected, e.g., over the course of multiple experiments. For example, the methods can be applied to a collection of decovoluted spectra from multiple variant sets that may include data from aglycosylated and glycosylated molecules. The described methods can be used to anchor these two classes and cluster the most similar spectra. From the aligned data from step 608, a second rectangular matrix is generated (304) that can be processed further as described below, e.g., using similarity metrics, correlative statistics (606) and / or visualization methods (107).
[0100] FIG.3B is a flow chart illustrating an iterative approach where, after data acquisition (101) and deconvolution (102), mass spectra / distributions may undergo global optimized alignment processing 305. In the context of FIG. 3B, the known or estimated mass of each analyte may be subtracted (104) from the corresponding measured mass of a mass spectrum / distribution. This may be followed by pairwise calculation (306) of mass distribution similarity using the adjusted mass distributions. A global mass axis, where each mass value (e.g., x-value) in the whole data set across spectra has a corresponding mass intensity value (e.g., y-value), can then be generated (104). Aligned mass spectra / distributions can then be generated (110) from using the adjusted mass spectra / distributions and the global mass axis. From the aligned data a first (e.g., rectangular) matrix may be generated (105). Some fraction of the samples / adjusted mass distributions are likely to be well correlated - these represent the samples / adjusted mass distributions that (a) have similar heterogeneity profiles, and (b) where the known difference (e.g., the expected protein sequence) is present in the sample. These correlated samples / aligned mass distributions may be identified (301) and / or removed from the matrix and may then be used as anchors against which all other samples / aligned mass distributions are aligned (302) by minimizing the residual difference between each poorly correlated sample / aligned mass distribution and the set of well correlated samples / aligned mass distributions (anchors). The mass difference that minimizes the residual difference between the anchor samples / aligned mass distributions and the sample / aligned mass distribution under investigation may be determined (303) to be the mass of the underlying molecule (e.g., protein). This determination is particularly strong if - after alignment - the similarity score (or correlation coefficient) is strong between the aligned mass distribution under consideration and the anchor set. Similar to the other approaches, this method is well suited for protein variant sets with sequences containing a limited number of amino acid substitutions on a relatively constant backbone and expressed in the same expression system. The approach described herein is also well suited for the alignment of dissimilar data sets where dissimilar molecules (e.g., proteins) may be present and / or the approach can be used for untargeted clustering and aggregation of all data collected, e.g., over the course of multiple experiments. For example, the methods can be applied to a collection of decovoluted mass spectra from multiple variant sets that may include data from aglycosylated and glycosylated molecules. The described methods can be used to anchor these two classes and cluster the most similar mass spectra. From the aligned data from step 302, a second (e.g., rectangular) matrix can be generated (304) that can be processed further as described herein with respect to step 112.Data Acquisition and Deconvolution
[0101] Raw data that can be used as input for the technologies described in this specification can be acquired using any number of different mass determining methods, including one or more mass spectrometry-based ionization techniques and / or mass analyzers and / or charge analyzers, as well as other techniques such as mass photometry and / or nanomechanical mass spectrometers, each alone or in combination as further described below. A requirement for the approach described herein is that the end data product results in a measurement of a sample’s underlying mass distribution.
[0102] FIG. 4 illustrates an example system according to an example implementation. The system 400 includes an analyzer 401, e.g., a mass spectrometer arranged to obtain raw mass spectrometry data from one or more samples (e.g., proteins). In some example implementations, multiple, e.g., biological, samples are arranged in a multi-well plate and analyzed with a 6545XT AdvanceBio LC / Q-TOF (Liquid Chromatography / Quadruple-Time of Flight) mass spectrometer, manufactured by Agilent®, using ESI (electrospray ionization) to generate ions. In other implementations, different mass spectrometers, mass spectrometry techniques, and / or ionization methods may be used. In this example, multiple, e.g., biological samples are arranged in a multi-well plate 402 and analyzed, e.g., simultaneously, to produce mass spectra.
[0103] The raw mass spectrometry data is received by a computer 403. In the example shown in FIG. 4 , the computer includes a processing arrangement, such as a central processing unit (CPU) 404 comprising one or more processors and / or microprocessors. The computer also includes one or more storages 405 for storing data and programs for execution by the processing arrangement, an input / output interface 406 for input and output of data, and a receiver 407 for receiving data, such as the raw data from the analyzer 401 (e.g., mass spectrometer).
[0104] The raw data, output from the analyzer 401, can be deconvoluted by the computer 403 to generate mass spectra and / or corresponding mass distributions for the different samples. The computer 403 can be used to perform the subsequent processing steps described in this specification, or the data can be transferred to one or more additional computers / computing systems for processing.
[0105] The types of data acquisition (e.g., mass spectrometry) methods used may dictate one or more required intermediate processing steps needed to arrive at the mass distributions required for mass distribution alignment as described in this specification. In someembodiments, one or more deconvolution processing may be performed. In the context of mass spectrometry, deconvolution refers to taking m / z spectra, where the m / z values are measured but the mass (m) and the charge (z) are unknown. Sets of m and z are iteratively explored such that the experimental mass spectrum (m / z spectrum) is maximally explained by the set of masses and charges. For example, electrospray ionization (e.g., high-flow pneumatically assisted electrospray, static nano electrospray, desorption electrospray ionization, etc.) utilizing a mass analyzer that produces the conventional mass spectrum (intensity as a function of m / z - where m / z is the mass-to-charge ratio) can produce a plurality of charge states and require deconvolution. In some implementations, electrospray ionization applied to charge detection mass spectrometry (CDMS) utilizing a mass analyzer (e.g., Orbitraps, Fourier transform ion cyclotron resonance, electrostatic linear ion traps, etc.) that measures both the m / z and the charge of each individual ion may not require deconvolution. CDMS may simply require multiplication of the m / z by the charge to arrive at the mass distribution although deconvolution algorithms can also be applied to these types of data (see, e.g., UnidecCD: https: / / doi.org / 10.1021 / acs.analchem. lc03181).
[0106] In some implementations, deconvolution processes as described in this specification are optional. In some implementations, mass spectrometry methods utilizing MALDI mass spectrometry, where singly charged ions are primarily generated and analyzed using mass analyzers, may be directly utilized without deconvolution since a mass spectrum where z = 1 reduces to the underlying mass distribution. In MALDI experiments, where the charge state distribution produces a plurality of charge states, deconvolution and other pre-processing steps may be required.
[0107] In some implementations, the methods described above can be used in combination with any number of “hyphenated techniques” for both online and offline analysis. Liquid chromatography mass spectrometry (LC-MS) can be used to introduce samples into the analyzer 401 utilizing either online or offline separations. Even higher throughput systems can greatly benefit from this approach - these systems include dedicated high-throughput systems like the Agilent® RapidFire and acoustic droplet ejection (EchoMS from Sciex®) technologies, which can generate data from thousands of individual samples.Mass Distribution Adjustment
[0108] Comparing mass distributions from samples (e.g., molecules, e.g., proteins) that differ in mass and may contain similar or differing heterogeneity profiles (or topology) is difficultboth visually and computationally since the heterogeneity profiles may exist along different regions of the mass distribution. FIG. 5A shows a graph of mass spectra with example mass distributions obtained from deconvolution of the received spectrometry data from two samples before mass adjustment (the table in FIG. 5B shows the original value matrix of the data of the graph in FIG. 5A ). Known / expected sample (e.g., molecule, e.g., protein) mass (7.3 for Sample 1 and 10 for Sample 2) was subtracted from each data point in each mass spectrum / distribution of FIGS. 5A and 5B, resulting in an alignment of the peaks / mass spectra as shown in FIG. 6 A (the table in FIG. 6B shows the adjusted value matrix of the data of the graph in FIG. 6A).
[0109] In the example of FIG. 6B, the measured values of Sample 1 and Sample 2 share a common mass axis (e.g., x-axis) but cannot be represented as a matrix without missing intensity values. Because the mass of Sample 1 and Sample 2 are known to a different number of decimal points (see FIG. 5B), the matrix of the adjusted mass distributions of Sample 1 and Sample 2 (FIG. 6B) contains “NA”-values (missing intensity values). In the example shown here, there are no shared intensity values along the adjusted mass axis between Sample 1 and Sample 2 after mass adjustment.
[0110] Using the adjusted mass distributions, a global mass axis, where each mass value (e.g., x-value) in the whole data set across spectra has a corresponding intensity value (e.g., y-value), may be generated. An example global mass axis can be generated by, e.g., (a) reducing the numerical accuracy of the “known” / “expected” masses of the samples (e.g., by reducing the known / expected masses of the samples by one or more decimal places, e.g., from 10.0 to 10 and 7.3 to 7 for Sample 1 and 2, respectively) or by globally adjusting mass values (e.g., x-values) of one sample to the nearest mass value (e.g., x-value) of another sample, (b) interpolating (using any art- and / or industry-known interpolation algorithm) missing intensity values for each sample to find intensity values (e.g., y-values) for all unique adjusted mass axis values contained within the set, or (c) interpolating missing intensity values for each sample to find intensity values (e.g., y-values) on an adjusted mass axis sampled at the resolution of the original mass axis (e.g., x-axis). With regard to option (c), rather than fill in all missing values in, e.g., FIG. 6B, which would sample the distributions at a frequency higher than the frequency at which they were originally sampled, one can define an x-axis range (xmin, xmax) and sampling frequency across that range (step size) and then interpolate at those points which have no value. The graph in FIG. 7A shows the example adjusted mass distributions from FIG. 5A and FIG. 6A aligned to a global mass axis (thetable in FIG. 7B shows the value matrix of the (globally) aligned mass distributions of the graph in FIG. 7A).
[0111] Aligned mass distributions can be normalized so their mass intensity values span a range (e.g., from 0 to 1). This may be known as “min-max normalization.” In some instances, an intensity value may be normalized using the following equation: intensity value— minimum intensity value ,- = normalized intensity value Equation 1 maximum intensity value -minimum intensity value
[0112] The aligned mass distributions (optionally normalized mass distributions) can be compared pairwise and / or groupwise using a variety of methods, e.g., Pearson’s r value, linear regression (slope, y-intercept), and / or spectral contrast angle. The matrix of (globally) aligned mass distributions (e.g., as shown in FIG. 7B) can be processed further as described below, e.g., using similarity metrics, correlative statistics, and / or visualization methods.
[0113] FIGS. 8A-8B show example mass distributions obtained from deconvolution of the received spectrometry data from a plurality of samples. FIG. 8A shows mass distributions for five protein samples. Because the spectra overlap one another, visual analysis and identification of features contributed by a particular component of the samples is not straightforward, even with this relatively small number of samples. As the number of samples increases, visual inspection and analysis becomes even more difficult, as shown by the mass distribution in FIG. 8B, which is based on 10 samples. In a similar way, identification and / or annotations of features in the mass distributions using a computer- implemented method also becomes more difficult as the number of samples increases. This task becomes nearly impossible when comparing hundreds or thousands of individual mass distributions. Several approaches to resolve this problem are described above in this specification, e.g., using the methods for alignment described herein.
[0114] FIG. 9A shows examples of adjusted mass distributions produced using methods as described herein for the mass distributions of FIG. 8A, so that the peaks and troughs of the mass distributions occur at the same adjusted mass. FIG. 9B shows an example of aligned mass distributions and signal intensity (y-axis) values resulting from the mass distribution alignment processing described herein being applied to the adjusted mass distributions of FIG. 9A. It is noted the aligned mass distributions of FIG. 9B have been normalized to a range of 0 to 1.
[0115] An outcome of these procedures is that each mass distribution in the set is aligned to allow for improved visual comparison. The mass distributions can be represented as a matrixamenable for a wide variety of statistical and machine learning-based techniques, e.g., as described in this specification.Modification Adjustment
[0116] Various chemical treatments or modifications can be applied to a material that can be analyzed using spectrometric techniques, e.g., oxidation, reduction, hydrogenation, sulfurization, and the like. Biomolecules, for example, proteins, can be modified to provide a number of (additional) options for analytical methods, e.g., through (de-)phosphorylation, (de-)glycosylation, (de-)ubiquitination, (de-)methylation, (de-)lipidation, or proteolysis. Each treatment or modification can be performed alone or in combination with any other treatment. A first treatment or modification can occur at the same time as a second treatment or modification, or two (or more) treatments can occur sequentially. In some implementations, a first treatment or modification is the same as the second treatment or modification (e.g., two or more phosphorylation steps can be performed sequentially). In some implementations, a first treatment or modification is different from the second treatment or modification (e.g., phosphorylation and glycosylation can be performed at the same time or can be performed sequentially). For example, a method to reduce heterogeneity in protein samples includes deglycosylation of proteins, e.g., using the enzyme PNGase F and / or other enzymes or enzyme mixes. PNGase F deglycosylates proteins with N-linked glycans by cutting between asparagine and carbohydrate chains. This reaction with PNGase F selectively cleaves N- linked glycans from the protein resulting in a “deglycosylated” protein sample with less heterogeneity compared to the untreated sample. The deglycosylated sample can then be analyzed by mass spectrometry. However, the deconvoluted mass distribution / spectrum of untreated and treated protein becomes difficult to directly compare with or without mass adjustment based on the theoretical protein mass. In both cases, the main peak of interest in each sample would exist at a different mass value or adjusted mass value along an axis (e.g., the x-axis) of the original and adjusted mass distribution, e.g., as illustrated in FIG. 10. In this example, the non-reduced, deglycosylated sample (having the leftmost most intense peak in FIG. 10) and the untreated sample (having the rightmost most intense peak in FIG. 10) were mass adjusted using the expected sequence mass of the Immunoglobulin G (IgG) antibody calculated from the stoichiometry of expected assembly (here, 2 light-chains and 2 heavy chains and adjusted for the expected number of disulfide bonds). The most intense peaks are positioned at two different values along the adjusted mass axis because they differby the mass of two glycans. Incomplete removal of glycans from the sample results in three distinct peaks (deglycosylated sample - loss of zero, one, or two glycans), one of which has the same mass as the original sample but at lower intensity.
[0117] This illustrated misalignment can be addressed by expanding the possible values by which a sample (e.g., protein) can be mass adjusted to include expected modifications. The inclusion of expected modifications or transformations can be utilized during or after mass adjustment. An example workflow is shown in FIG. 11 A, which is based on the workflow in FIG. 1A. In the modified workflow, mass adjustment is performed through a mass subtraction step (103a) and a modification adjustment step (103b). Another example workflow is shown in FIG. 11B, which is based on the workflow in FIG. IB. In the modified workflow, global optimized alignment (1101) includes mass adjustment (i.e., generation of the adjusted mass spectra / distributions) being performed through the subtraction step (109a) (i.e., where the known / expected sample mass is subtracted from the measured sample mass) and a modification adjustment step (109b), which adjusts the mass (e.g., x) axis of a mass distribution based on what is known about a sample (e.g., the sample’s mass, chemical modifications applied to the sample, a combination of both, etc.). Assuming two samples - A and B - which have known masses - A mass and B mass -the x-axis in the spectra of A and B can be adjusted - A specfx] and B specfx] - such that the two samples are aligned by the following:A_aligned[x] = A_spec[x] - A_mass B alignedfx] = B specfx] - B massA specfx] and B specfx] are the original, unmodified x-axis values of A and B.A alignedfx] and B alignedfx] are the adjusted values of the x-axis for A and B. Subtracting the known masses of A and B from their respective deconvoluted spectra (A specfx] and B specfx]) would result in the alignment of the two different samples.
[0118] if A is modified by some chemical method (e.g., enzymatic method, etc.), the mass can be added or subtracted from the modified / unmodified sample such that the deconvoluted spectra of A (unmodified) and A (modified) are aligned:A_mass = A_amino_acid_sequence_mass + glycan_massA mass modified = amino acid sequence mass A (the expected mass of A after modification is the sequence mass - the modification process would be deglycosylation with enzymes)
[0119] The spectra of unmodified sample A can be adjusted by subtracting:A_spec[x] - glycan mass (subtract the mass of the glycan present in unmodified A)
[0120] This results in the spectrum of unmodified A directly aligning with the spectrum of modified A.
[0121] Multiple options exist to manipulate masses. Using the example above, it is known unmodified sample A and modified sample A are composed as follows: A_mass = A_amino_acid_sequence_mass + glycan_mass A mass modified = amino acid sequence mass A
[0122] Both the modified and unmodified spectra can be adjusted using the following: A_aligned_unmodified[x] = A_spec_unmodified[x] - A amino acid sequence mass- glycan_massA_aligned_modified[x] = A_spec_modified[x] - A amino acid sequence mass
[0123] Now, both the modified and unmodified spectra align. This can be extended to deal with a different unmodified sample, B:B_aligned_unmodified[x] = A_spec_unmodified[x] - B amino acid sequence mass- glycan mass
[0124] Since all known mass components of all samples have been removed, A unmodified, A modified, and B unmodified are all aligned.
[0125] In an example implementation, a deglycosylated sample can be mass adjusted by subtracting the mass of the IgG minus the mass of 2x glycans such that the most intense peak is expected to land on the most intense peak in the untreated sample, which is assumed to have, a priori, a specific modification profile. This approach assumes a priori a specific chemical transformation will occur. FIGS. 12A and 12B illustrate mass adjustment including the adjustment for the modification. The mass adjustment of the treated (deglycosylated) sample includes adjustment for the removal of N-linked glycans resulting in perfect alignment with the untreated sample based on information about the chemical transformations performed on the sample.
[0126] The molecule and expected modification profile can be subtracted at various points during the analysis process, e.g., at the start or after globally optimized alignment. In some implementations, the molecule and expected modification profile can be subtracted at the start of each analysis. For example, the untreated IgG may be mass adjusted by accounting for common, expected modifications (e.g., loss of C-terminal lysine, the formation of pyroglutamate, and / or common glycans). The treated sample can then be mass adjusted by accounting for the same set of modifications except for the glycans which have already beenremoved (the modification set would only include loss of C-terminal lysine and the formation of pyroglutamate). This method would position the adjusted mass distributions at a value of zero if the molecule (e.g., protein) had the expected composition (e.g., sequence) and modification profile. This approach assumes a priori a specific set of modifications exist for a certain molecule or molecule set.
[0127] The methods for modification adjustment described above can be used with any type of sample (e.g., biological sample, e.g., protein) modification or chemical heterogeneity of the analyte or the reference sample (e.g., biological sample, e.g., protein) including, but not limited to, glycosylation, phosphorylation, acetylation, methylation, ubiquitination, sumoylation, lipidation, proteolysis, deamidation, hydroxylation, and / or molecule (e.g., protein) labeling.Generation of the Matrix
[0128] As described above, in the first step mass distributions may either be computed or directly measured. Second, the set of mass distributions may be adjusted and aligned (e.g., using either subtraction of known differences, unknown global alignment, or hybrid iterative methods using anchored alignment) into a matrix with a shared, global mass axis.
[0129] When analyzing multiple data sets, it is possible that the mass axis of each adjusted mass distribution does not have the same values as the other axes across the set. In some implementations, a set of mass distributions could be sampled at integer mass resolution. If Mcaic is determined with a resolution of 0.1 Da, subtraction of Mdist - Mcaic would result in floating point mass values not originally present in the mass distribution. In the same set, another value of Mcaic corresponding to a different molecule may have been calculated with the same decimal point resolution but land on an integer value. Now, two sets of mass values exist in the set of adjusted mass distributions, and this problem becomes more pronounced if mass distributions or expected masses are represented with an increasingly large number of significant figures. Thus, the calculation of a global mass axis for the set of adjusted mass distributions is beneficial for the generation of a matrix upon which further processing (e.g., similarity, correlative, etc.) can be performed as described herein.
[0130] In some implementations, the procedure is configured to ensure that (1) Mdist is sampled at the same interval across the set of samples and / or (2) Mcaic is calculated with the same numerical accuracy as Mdist. For all other cases, other methods may be implemented for generating a global mass axis.
[0131] In some implementations, mass binning can be used. In mass binning, the mass distribution with the lowest resolution mass axis (e.g., x-axis) can be identified within a set and used to define the bounds of each mass bin. Every other mass distribution may then be binned by collapsing the higher-numerical accuracy mass distributions into the global, low resolution mass axis.
[0132] In some implementations, interpolation methods can be used. In interpolation methods, the opposite approach can be taken. Here, a higher numerical accuracy mass axis is defined. Two example methods include: (I) Unique value set - find all unique mass values across the set, sort the mass values in increasing order, and use this order as the global mass axis, and (II) Highest numerical accuracy - select the highest numerical accuracy mass axis.
[0133] Both approaches (I and II) for selecting a global mass axis potentially produce sparse matrices and, in the extreme case, can result in a situation where two mass distributions have no shared intensity values. This potential problem is solved by implementing interpolation methods to fill missing values within the sparse matrix. Various methods exist and can be applied depending on the specific data set used. Examples for interpolation methods include the art- and industry-known techniques of bilinear and bicubic spline interpolation of irregular data.
[0134] Following any mass binning or interpolations operations, aligned mass spectra / distributions may be generated by gridding the set of adjusted mass distributions into a matrix with a shared, global mass axis. The establishment of a matrix having a global mass axis may be followed by the calculation of similarity metrics as described in this specification.
[0135] FIG. 13 illustrates an example analysis workflow from mass-aligned deconvoluted spectra to calculation of pairwise statistics to conversion to a pairwise square matrix. The calculation of pairwise similarity metrics does not require generation of a matrix containing the entire set of adjusted mass distributions and can be performed after generation of the global mass axis. Any of the methods described above with respect to the generation of the global mass axis and matrix of aligned mass distributions can be utilized for pairwise comparison of mass distributions, which may simplify processing of data. After alignment, two mass distributions can be directly overlaid, plotted as a mirror plot, where one sample is plotted in the positive-axis (e.g., positive-y) direction and the other is plotted in the negativeaxis (negative-y) direction, or the aligned pairwise intensity values can be plotted (see FIG. 13, middle panel). This approach provides detailed comparison of two samples that more readily shows differences in the mass distribution compared to traditional methods where themass distributions may exist along disparate regions of the mass axis (e.g., x-axis). Similarity metrics can be calculated using any number of methods once the adjusted mass distributions are gridded onto a shared, global mass axis (i.e., are represented as aligned mass distributions). Examples of similarity metrics include Pearson’s r, r2value, cosine similarity, and / or spectral contrast angle.
[0136] Groupwise comparisons can also be performed. Groups, clusters, and related spectra can be identified using manual, unsupervised, and / or supervised methods. The mass distributions of the members of the group can be combined or averaged to generate a “consensus” mass distribution. Then, groups, clusters, and related spectra can be compared using binary comparisons (see above) reducing the number of spectra that need to be compared. For example, a group of 100 similar mass distributions could be used to generate a consensus mass distribution. Another set of 100 related mass distributions (but differing from the first group) can be used to generate a second consensus mass distribution. These consensus mass distributions can be compared using binary comparison resulting in a single figure representing information from 200 individual mass distributions that, if compared combinatorially, would result in thousands of individual figures.
[0137] The generation of the matrix of aligned mass distributions can be followed by a vast number of statistical approaches to be applied to the data for unsupervised interpretation of the data. These approaches can be applied to the (e.g., rectangular matrix) of aligned mass distributions or to a (e.g., square) matrix resulting from pairwise similarity calculations.
[0138] In some implementations, unsupervised methods can be applied to the (e.g., rectangular) matrix of aligned mass distributions. In some implementations, unsupervised methods can be applied to the (e.g., square) matrix of pairwise similarities.
[0139] Positive matrix factorization (or non-negative matrix factorization) and principal component analysis are attractive methods for dimensionality reduction and clustering of the aligned mass distributions. Positive matrix factorization has the added benefit of producing factors that are human interpretable.
[0140] Visualization and ordering of the (e.g., square) matrix of pairwise similarities via a seriation method can provide visualization of related or similar mass-adjusted distributions. Unsupervised clustering can also be performed using a variety of algorithms - K-means, Density -based spatial clustering of applications with noise (DBSCAN), Hierarchical Density- Based Spatial Clustering of Applications with Noise (HDBSCAN), Uniform Manifold Approximation and Projection (UMAP), t-distributed stochastic neighbor embedding (t-SNE), etc. Example visualization workflows from (e.g., square) matrices of pairwise similarities to seriation to clustering are illustrated in FIGS. 14A and 14B. The top images in FIGS. 14A and 14B are un-ordered matrices, which are transformed into ordered matrices (middle images in FIGS. 14A and 14B). General trends, groups, and overall relatedness of the whole data set can be visualized in a single, scalable figure. Fundamentally, these visualizations consist of some pairwise measure of similarity for each mass distribution within a set against all other mass distributions within a set (bottom images in FIGS. 14A and 14B). Further, visualization with network graphs and dendrograms can also highlight underlying structure and identify groups of related / unrelated spectra.
[0141] Distinct groups of mass distributions can be assigned using any number of widely available clustering algorithms based on the pairwise similarity or correlation between all members of the set. The membership and relatedness of each cluster can be visually evaluated using dendrograms and network graphs. Dimensionality reduced visualization, which includes positive matrix factorization (non-negative matrix factorization) applied to a set of mass distributions, can produce factors which represent groups or underlying features. These factors can be utilized directly as human interpretable mass distributions. These visualizations are provided by generating the adjusted mass distributions of the present disclosure. Similarly, principal component analysis can be utilized. In both cases, grouping can be determined by examining loadings. Principal components are generally less interpretable compared to the factors generated by positive matrix factorization. tSNE and other methods offer similar approaches to grouping data and identifying underlying trends present within large sets of mass distributions.Grouping / Clustering
[0142] The technologies described in this specification can be used in concert with traditional (manual) annotation / verification approaches or used with non-targeted analysis techniques. While such non-targeted techniques may still require some a priori knowledge of molecule mass, annotations or other knowledge regarding any modifications of the analyte are not required.
[0143] High-throughput data generation can eventually reach the scale where manual quality control, let alone analysis, becomes virtually impossible. For high-throughput mass spectrometry experiments, mass / modification adjustment as described herein followed by unsupervised clustering methods provides a route to data reduction that can provide high-quality data generation at scale. Described in this specification are example clustering methods based on mass adjusted deconvoluted spectra (mass histograms) using liquid chromatography mass spectrometry for a set of molecule variants. The data was mass adjusted and a pairwise correlation matrix generated as described in this specification. Pairwise results are shown in FIG. 15A.
[0144] The correlation matrix was then fed into a density-based spatial clustering of applications with noise (DBSCAN) algorithm for cluster assignment based on the correlation coefficient values defined as: distance metric = 1 - r2. In an example, DBSCAN was run with epsilon value of 0.01, where epsilon is a parameter specifying the radius of a neighborhood with respect to some point. This parameter can be adjusted to generate more or fewer clusters. For example, Epsilon values can be between 0.0001 and 0.99, between 0.001 and 0.1, between 0.002 and 0.05, or between 0.005 and 0.01. In some implementations, clustering is performed using two or more Epsilon values (e.g., 2, 3, 4, 5, 10 or more) either iteratively or in parallel to evaluate cluster stability. Adjusting the Epsilon values can result in, e.g., between 1 and 100 clusters, between 2 and 20 clusters, or between 3 and 10 clusters, e.g., 3, 4, 5, 6, 7, 8, 9, or 10 clusters. The result can be visualized by plotting the samples using, e.g., the same order used for visualizing the pairwise correlation matrix. In some implementations, the correlation can be re-ordered / re-arranged using the DBSCAN clustering results to help facilitate visual inspection. Each pairwise cell can be assigned a value of -1 if their cluster group assignment differs. Otherwise, each pairwise cell can be assigned the value of the cluster group if both samples are assigned to the same cluster.
[0145] Lastly, the adjusted mass distributions within a single cluster group can be averaged to generate an average adjusted mass distribution. In this example, the selection of an epsilon value of 0.01 ensures that the samples contained in a cluster group are extremely similar. The clustering results are shown in FIG. 15B. The averaged adjusted mass distributions can be manually evaluated to better understand what these clusters represent chemically. For example, the averaged mass distributions can be manually or automatically annotated, e.g., using a conventional annotation procedure, and all members of the cluster group could be assigned a mass verification status based on the agreement between the annotation and the peak location. Such an approach can provide high-quality data triage at scale. In this example, 179 samples were reduced to eight clusters where one of the clusters represents global outliers and five clusters represent clusters with highly similar samples.
[0146] In parallel to this procedure, samples were analyzed using a conventional annotation procedure to investigate how well annotation at the cluster group level would compare to conventional analysis. The two largest cluster groups (three and four) contained 31 and 76 samples, respectively. Every member of these clusters was mass verified by conventional analysis, which translated to almost 60% of the samples being mass verified without ever considering the original data. Cluster groups two, six, and seven contained no members which were mass verified by the conventional method. Cluster group one contained only four samples, but all samples were mass verified. This raised the number of samples which belong to unambiguous clusters to almost 70%. The only cluster group with a mixed population was cluster group five and it contained only seven samples. Lastly, 43 samples were identified as noise with -42% being unverified.
[0147] With true high-throughput data including, e.g., several thousand samples, reliance on the clustering methods described in this specification may be necessary. For moderately sized data (100-1,000 samples), a combined approach of clustering and conventional analysis can be synergistic. This analysis raises questions about the 43 samples that are not highly correlated with any of the other samples, and further investigation could be conducted. These approaches guide the analyst to understand where their efforts should be spent analyzing the complexities of these data sets and can aid uncovering unexpected, subtle, or even egregious deviations from desired behavior.Example
[0148] The grouping technologies were evaluated using simulated IgG antibody spectra. A peak shape model was defined that was representative of IgG antibodies that could be used to generate realistic deconvoluted spectra. A set of simulated antibodies were created based on the sequence of IgG, which had different masses calculated from a random set of point mutations. Random mutations included mutations of between 1 and 20 residues in the heavy chain, where location for each mutation was randomly selected. Moreover, post-translational modifications were introduced. Six sets of modified antibodies were created as shown in the Table 1 below:TABLE 1
[0149] Where “Lys-Loss” denotes Lysine clipping, “PyroGlu” denotes Pyroglutamic acid, G# refers to the number of Galactose on the two arms, “T=>V” denotes Threonine replacement with Valine, “M => Q” denotes Methionine replacement with glutamine, and “R => Y” denotes Arginine replacement with Tyrosine. Random noise (jitter) was added to the peak location and amplitude. Additional random noise was added across each spectrum, and the data was smoothed and normalized.
[0150] FIG. 16 shows 150 deconvoluted spectra of the modified simulated IgG spectra. Individual peaks are virtually undiscernible by the naked eye. FIG. 17A shows the 150 deconvoluted IgG spectra after mass adjustment. The spectra are well aligned with four peaks that appear virtually identical to the naked eye. Only Mod Set 2 exhibits a fifth peak (arrow), as further illustrated in FIG. 17B, showing aligned spectra of Mod Sets 1, 2, and 3.
[0151] FIG. 18A shows the pairwise correlation matrix with markings for the six mod sets. These markings are based on ground truth information illustrated in FIG. 18B. Based on visual inspection of the pairwise correlation matrix, Mod Set 2 can be clearly distinguished, while Mod Sets 1, 4, and 5 appear almost indistinguishable (similarly, Mod Sets 3 and 6 appear to belong to the same set).
[0152] The DBSCAN clustering using a similarity threshold of epsilon=0.01 (FIG. 19A) results in a grouping similar to visual inspection. Changing the epsilon value to 0.005 and thus increasing the sensitivity of the model results in clustering that closely represents the six Mod Sets as shown in FIG. 19B.
[0153] The simulation results show that correlation techniques plus clustering (e.g., DBSCAN) can provide grouping of (extremely) similar spectra and differentiation of subtle differences. In the example above, Mod Sets 1 and 3 are visually indistinguishable as shown in FIG. 17B yet can be readily differentiated with correlation and clustering analysis. Distinguishing features, e.g., the presence of additional glycan, are both visually identifiable after mass adjustment and easily identified with correlation and clustering. Subtle mass shifts are more difficult to distinguish at the simulated instrument accuracy. Decreasing DBSCAN epsilon neighborhood value allows these groups to be identified.
[0154] Clustering as described in this specification can increase the efficiency of analytical processes, e.g., for protein quality control. For example, undesired and / or unexpected modifications can be efficiently identified: only one sample of potentially hundreds ofsamples need to be tested and characterized because all members of the same cluster will (most likely) share the same modification.Machine Learning Implementation
[0155] Described in this specification are methods that provide untargeted analysis, clustering, and comparison of mass distributions generated from various analytical instruments. The alignment processes described in this specification can be used with unsupervised machine learning methods for classification or clustering of mass distributions at scale, to train an ML clustering / grouping model. For example, the generation of cluster groups from individual sample sets can provide labels for supervised machine learning with the goal of assigning group membership directly from mass distributions with or without knowledge of molecular weight, molecular sequence (e.g., DNA, RNA, protein, etc.), or structure (predicted or experimental structure information). The output of the ML model can be, e.g., alignment information and / or cluster ID. In some implementations, the models / model outputs can include human annotations of identified clusters and / or the results of the traditional annotation workflow for each sample or cluster (verified / unverified and / or annotations). FIGS. 20 and 21 illustrate an example implementation of the whole workflow whereby sets of samples are mass analyzed and data accumulated over time. The mass distributions are aligned and clustered during the analysis of each data set, e.g., as described above. These results from the clustering process (2101) are stored in an accessible database. At the time of training the model (2102), the data are accessed and mass distributions along with their assigned cluster identifier (and sequence or molecular weight) are provided to the model. Alternatively or additionally, all adjusted mass distributions can be re-analyzed en masse providing a globally consistent set of cluster labels.
[0156] Training of the proposed model can eliminate the discrete operations of alignment, statistical interpretation, and clustering whereby a set of data enters the model and group assignments are generated, as illustrated in FIGS. 22 and 23. A model 2301 trained using the cluster membership and raw data as illustrated in FIGS. 20 and 21 may optionally use the output (attributes) derived from model 2102 to perform grouping / clustering of the mass distributions. Grouping / clustering of aligned mass distributions may be an optional step. Further additions to the model can include annotations of individual features within each cluster (e.g., the set of post-translational and chemical modifications shared by the cluster) or human derived descriptions of the information contained in the cluster. For the latterexample, human derived descriptions can be accumulated over time as data are generated. These additional labels can be further engineered or summarized en masse (e.g., by using large language models) prior to being fed into the model for training. These approaches can bridge the non-targeted, annotation-free mass adjustment workflow and conventional annotation-based workflows.Applications
[0157] Applications of the technologies described in this specification include, e.g., (1) comparison of mass-variant species distributions reflective of molecular attribute distributions between biobetters (e.g., antibodies that target the same validated epitope as a marketed antibody, but have been engineered to have improved properties) and a reference with the aim to test the extent of similarities in such distributions of biobetters to the reference, (2) comparison of said distributions between biosimilars and references with the aim to test the extent of said similarities of biosimilars to the reference, (3) comparison of said distributions between biomolecule purifications from different cell lines with the aim to contribute to cell-line selection, (4) comparison of said distributions between biomolecule purifications from different manufacture processes with the aim to contribute to process development, and / or (5) quality control of high-throughput intact mass analysis or quality control for the verification of intended sequence using measurement of mass distributions, and others.
[0158] Mass spectrometry is well suited for the in-depth analysis of individual molecules. Current technologies for high-throughput molecule analysis typically lack approaches for dealing with hundreds or thousands of individual data sets / files. The need for high-volume, high-speed mass spectrometry analysis remains an under-addressed as data generation vastly outpaces the ability to review and ultimately act on these data.
[0159] Broad screening of molecules by intact mass analysis, e.g., using mass spectrometry, results in data that contain the full underlying mass distribution of the molecule sample (the deconvoluted spectrum). The heterogeneity of a molecular sample is measured as the deconvoluted spectrum. Describing and annotating these complex deconvoluted spectra is challenging - particularly in the absence of supporting experimental data (e.g., peptide mapping and released glycan analysis). Without orthogonal supporting data, a search space used to describe the mass distribution (deconvoluted spectrum) is poorly constrained. Best guesses based on frequently observed post-translational modifications are often used and arebased on previously observed post-translational modifications from the same or similar expression systems, proteins modalities, expression conditions, etc.
[0160] In the technologies described in this specification, the search space is defined as the allowed post-translational modifications that can be combinatorically applied to adjust the mass of the expected sequence. As search space grows, the number of false annotations increases and the ability to annotate spectra with high confidence decreases.
[0161] The technologies described in this specification can be used in concert with traditional annotation / verification approaches or applied independently as a non-targeted consensusbased method. This method utilizes all available data rather than confirming presence or absence of a particular molecule based on incomplete data utilization.
[0162] The technologies described in this specification readily provide visual comparison of deconvoluted data, which can be otherwise challenging or impossible when the non-adjusted deconvoluted mass distributions are overlaid.
[0163] The technologies described in this specification can be used to support the generation of consensus modification profiles (the search space). Deconvoluted spectra that cluster together (are highly correlated) can be assigned the same annotations. Large numbers of clustered spectra can be used to define the search space for all other molecules in the set or define the search space for all molecules of a particular modality for a given expression system (cell type) or expression platform. An expression platform can be or can include hardware systems including, but not limited to, bioreactor, plate, tube, and / or flasks.
[0164] The technologies described in this specification can generate consensus deconvoluted spectra, which can be generated by averaging multiple deconvoluted spectra together based on the clustering (grouping) results.
[0165] The technologies described in this specification are at least in part based on the assumption that an expected sequence mass is present in a sample. If this assumption is incorrect, then the alignment of dominant features arising from the various chemical modifications of the molecule will not occur and poor correlation will be observed.
[0166] Anomalies that would be (completely) missed in traditional annotation schemes can be readily identified using the technologies described in this specification. For example, the presence of additional glycans at high mass, but with moderate abundance, are generally not considered in traditional approaches if the most abundant peak in the deconvoluted spectrum is annotated. These high-mass signals can impact similarity measures and lower Pearson’s correlation coefficient or spectra similarity.
[0167] The technologies described in this specification are well suited for high-throughput non-deglycosylated intact mass analysis, but can also be applied to any molecule (e.g., protein) analysis that produces mass spectra that require deconvolution, particularly methods that are run repetitiously (e.g., high-throughput). For example, the technologies can be used for analysis of treatment combinations, e.g.: non-reduced non-deglycosylated, non-reduced deglycosylated, reduced non-deglycosylated, and / or reduced deglycosylated. As described above, the a priori knowledge of the system can be used to perform mass adjustment for each of these samples based on what is known about them individually, but the mass distributions can be adjusted and overlaid on the same global axis for analysis. Other examples include reduced non-deglycosylated antibody analysis, whereby the analysis could be performed on deconvoluted spectra arising from a set of heavy chains and a set of light chains and / or include deglycosylated antibody analysis.
[0168] Applications of the technologies described in this specification are not limited to liquid chromatography-based mass spectrometry. They are equally applicable, e.g., to other techniques of introducing molecule (e.g., protein) into mass spectrometers. These techniques include flow injection systems (liquid chromatograph without a column consisting of pumps and autosampler, dedicated high-throughput systems), manual and automated nanospray systems (pulled glass capillary), acoustic droplet ejection, desorption electrospray (DESI), or nano desorption electrospray (nDESI).
[0169] The technologies described in this specification are also applicable to mass spectrometry methods that do not require deconvolution, e.g., Matrix Assisted Laser Desorption / Ionization (MALDI) (which would not require deconvolution due to different ionization mechanism compared to electrospray) or charge detection mass spectrometry (CDMS), which directly measures the m / z and charge of individual ions simultaneously, thus providing direct calculation of mass without deconvolution. In the case of MALDI, mass spectra would be correlated / compared and in the case of CDMS the true mass spectrum (mass (Da) vs intensity) would be used in place of the deconvoluted mass spectrum.
[0170] The technologies described in this specification are also applicable to methods not based in traditional mass spectrometry. For example, mass photometry is capable of measuring mass distributions of native molecules (e.g., proteins) and molecule (e.g., protein) assemblies based on light scattering and interferometry. Emerging technologies, such as nano electro-mechanical system mass spectrometry, also produce data amenable to these approaches.
[0171] The techniques described above are thus applicable to the unsupervised analysis of high-throughput mass distributions acquired by any technique. This data-driven approach is advantageous because it enables biology-independent and chemistry -independent comparison across a data set or multiple sets. The alternative is to use “expert system” algorithms that depend on biological / chemical knowledge, are application-dependent, and whose usefulness is context-dependent, which limits flexibility and only solves the problem(s) the algorithm was intended to solve. In the technologies described in this specification, totally unsupervised analysis can be readily conducted by visualizing an ordered distance or similarity matrix, dendrograms, and / or network graphs. These techniques provide the use of computational tools for visualizing of the underlying structure in the data.
[0172] Workflows downstream from the approach as described above are tailored to their specific data, application, and / or scientific question. One example is high-throughput molecule (e.g., protein) quality control (QC), where intact molecules (e.g., proteins) are analyzed with mass spectrometry. Here, the goal is to determine if the intended amino acid sequence was expressed. Conventional annotation workflows are improved and made more robust by using the methods described in this specification because numerous artifacts or inconsistencies can be present in these data sets, especially in a high-throughput screening setting. Low intensity samples with low-signal to noise ratios may be annotated and “sequence verified” by conventional approaches but would be less correlated with well behaved, high-intensity samples. Such low intensity samples can be identified with the methods described in this specification.
[0173] The technologies described in this specification can be used for high-throughput oligonucleotide (e.g., DNA, RNA, or mRNA) quality control. Sufficiently large nucleotide- based polymers may require deconvolution, or analysis with charge detection mass spectrometry, to obtain a mass distribution. Here, the technologies described in this specification can be used to verify the sequence mass, verify paired / unpaired ratios or presence, and / or identify heterogeneity present in the sample. Comparison of many different (oligonucleotide) samples poses similar challenges as high-throughput molecule (e.g., protein) quality control whereby conventional analysis requires inspecting each sample’s file / spectrum / mass distribution individually. Mass adjustment provides a method for rapidly investigating the overall behavior of the samples and provides understanding from a single figure (adjusted mass distribution or correlation matrix) otherwise unobtainable through traditional techniques.
[0174] The technologies described in this specification can be used for the discovery of other complexities that would otherwise require manual data review. The presence of unexpected glycans (e.g., O-linked glycosylation patterns), unexpected cleavage events, and / or other modifications can produce multimodal mass distributions. A mass distribution’s most abundant peaks may be annotated successfully, but the presence of additional, unexpected features may not necessarily be annotated, let alone flagged for review, unless specific algorithms are implemented. The technologies described in this specification provide flexible and biology-independent methods for discovering features that may be present in only a small fraction of samples analyzed. Conventional analysis would require manual review and a trained analyst to find these features among hundreds or even thousands of individual mass distributions.
[0175] In some implementations, the technologies described in this specification can be used for the generation of consensus spectra. The techniques of the approach described above generate a set of adjusted mass distributions. Many high-throughput screening applications are focused on exploring combinations of a chemically complex space with large numbers of permutations. In the context of protein-based biomolecules, a protein sequence space can be explored for the optimization of a known binder or protein of interest. These molecules are frequently built upon relatively fixed backbones. Antibody exploration campaigns would be constructed from a single isotype, a single subclass within an isotype, and variants may be generated targeting only specific subregions or specific amino acids.
[0176] A collection of well-behaved adjusted mass distributions can then be averaged and used as a consensus or reference mass distribution. In some implementations, it may be expected that this mass distribution would be shifted along the mass axis depending on the exact molecular weight. This task could be automated by assuming that within a set of mass distributions most of the data are well-behaved. Then, the largest group of related spectra can be grouped, averaged, and assigned as a consensus spectrum representative of the high- throughput generation experiment. In some implementations, simple averaging along the global mass axis can be used. In some implementations, a positive matrix factorization can be used whereby all adjusted mass distributions, regardless of similarity or quality, can be included in the analysis. A factor representing the consensus mass distribution among a large number of spectra can then be selected.
[0177] These consensus spectra can be applied within the context of a single experiment or across experiments with varying degrees of similarity. For example, a high-throughputexperiment may have various attributes related to the production of the test materials: cell type used to express the protein, DNA plasmid backbone or transfection strategy, physical expression system (bioreactor, plate based, shaker flask, etc.), and / or conditions (or real-time monitoring) within the physical expression system. These experimental attributes may be related to other high-throughput generation experiments and can serve as a reference spectrum across projects / experiments that may or may not be biologically related. Moreover, these consensus spectra can represent the general behavior of a protein across high- throughput experiments or campaigns distilling possibly thousands of mass distributions for a single experiment to a single spectrum for QC across experiments.
[0178] The exact implementation of the consensus spectrum is application dependent. What data should be grouped, how those decisions should be made, etc. is context dependent. The techniques of the approach described above lay the foundation providing these techniques in a generalizable manner.
[0179] In some implementations, the technologies described in this specification can be used for re-annotation of mass distributions. Mass distributions with a high degree of similarity are expected to be comprised of highly conserved sets of chemical entities. Conventional molecule (e.g., protein) annotation workflows attempt to assign identity to peaks in the mass distribution based on the observed mass and an allowed set of modifications. The expected sequence and the modification set defines the search space against which the experimentally observed features in the mass distribution will be annotated. As with any algorithm, errors are likely to occur and incorrect annotations may be assigned due to subtle fluctuations in noise, peak location, peak shape, and / or the algorithm’s ability to determine peak location. This becomes especially problematic when using large search spaces (e.g., single protein sequence with many possible modifications).
[0180] The techniques described above followed by grouping / clustering of similar spectra opens an opportunity for re-annotating mass distributions which exhibit a high degree of similarity. If a feature within a set of similar, adjusted mass distributions are annotated with the majority of the annotations being the same, it may be justified to re-assign the annotations that differ from the dominant annotation. This approach can be thought of as occurring in parallel to the consensus spectrum generation described above whereby a consensus set of annotations is generated. The evidence for this reassignment is derived from the data as a whole rather than each mass distribution being evaluated in isolation.
[0181] In some implementations, the technologies described in this specification can be used for the analysis of biobetters or biosimilars.
[0182] Biobetters and biosimilars likely differ in mass or mass distribution compared to the original biomolecule (e.g., protein). The mass adjustment / alignment procedures described in this specification can provide direct comparison of a set of candidate biobetters and biosimilars against the original biomolecule (e.g., protein) and / or amongst themselves. Detection of unexpected post-translational modification is provided through mass adjustment / alignment followed by unsupervised statistical analysis. These methods can help guide the selection and development process towards candidate biomolecules (e.g., proteins) that have mass distribution profiles similar to the original biomolecule (e.g., protein).
[0183] In some implementations, the technologies described in this specification can be used for molecule (e.g., protein) therapeutic manufacturing, cell line development, and / or bioreactor optimization. The technologies described in this specification can provide, e.g., validation and / or optimization of manufacturing processes, e.g., by providing comparison of a material produced from one or more manufacture batches against a reference or set of reference, by providing comparison of all batches against all batches in a manufacturing run, or by providing comparison of a (candidate) biotheraputic molecule against other biotheraputics. Moreover, conclusions about the similarity of one sample to another can be made without conventional peak annotation. Additional insights can be gained by exploiting the adjusted mass distribution because many chemical heterogeneities are unresolvable at the intact level, and thus, would not be annotated by conventional methods. Unresolved heterogeneities at the intact level typically require laborious and time-consuming peptide mapping and multi-attribute monitoring for quantification. The technologies described in this specification provide an orthogonal approach with simplified sample preparation (e.g., no proteolytic digestion), no assumptions about chemical modifications, and nearly instantaneous analysis.
[0184] Early protein discovery campaigns generate a large number of sequence variants that can be analyzed by mass spectrometry (or related methods) as described above. For example, comparing the sequence-variant attributes using the technologies described in this specification, cell culture conditions can be evaluated between manufacturing batches of a given biopharmaceutical molecule (or set of sequence variants) providing more rapid control strategies compared to peptide mapping and multi-attribute monitoring. Late-stage biopharmaceutical development proceeds after selection of a lead molecule limiting the scopeof the problem to a single protein sequence generated under numerous conditions. These conditions can include the optimization and selection of manufacturing methods, selection of bioreactor conditions (or adjusting bioreactor conditions during online monitoring), and / or cell line selection. Example optimization methods include comparing the sequence-variant attributes of a candidate molecule against a validated / mature molecule with different molar mass (e.g., having a different sequence) using the technologies described in this specification. One major advantage of these techniques is that evaluating similarity of one sample to another is done independent of a user’s ability to annotate the data.
[0185] It may be expected that because a clonal protein sample is expressed, the data sets are well aligned without using mass distribution alignment. Regardless, alignment can still be conducted to provide a mass axis that represents the mass difference from the sequence mass. Then, the same rectangular matrix can be utilized for statistical analysis, the generation of consensus spectra, and / or for unsupervised analysis of the data. Alternatively, one could align the spectra using unknown global alignment. Well-behaved samples would all be shifted by the same value (essentially the sequence mass + / - the predominant modification profile) while samples exhibiting unexpected or differing modifications would be shifted by a different mass value.
[0186] In some implementations, the technologies described in this specification can be used for a number of research and machine-learning applications. Prediction of protein heterogeneity, O-linked glycans, the presence of non-consensus N-linked glycans, and other post-translational modifications are dependent upon a plurality of parameters surrounding the expression of protein based biotherapeutics. The measurement of the mass distribution of a sample is an appealing and information-rich opportunity for learning generalized principles between, e.g., sequence, chemical heterogeneity, other parameters related to process (bioreactors, plates, incubator conditions, etc.), and the expression system (cell type, expression system (e.g., cell free), etc.).
[0187] The mass adjustment / alignment procedure is the entry point to utilizing mass distributions directly without placing other algorithms between the mass distribution data and deep learning models. A conventional approach (i.e., no mass adjustment / alignment) applied to deconvoluted mass spectrometry data would require annotation of features within the mass distribution. These features could then be constructed into encodings amenable for machine learning / deep learning. This approach is problematic because (1) not all features will be annotated, (2) the search space must be known a priori, (3) redundant or ambiguousassignments will be present, and (4) the algorithms required for this step are application dependent and would need to be developed for each application rather than directly using the mass distribution. Mass adjustment / alignment using the mass distribution directly eliminates all these issues at the expense of a much larger matrix of mass distributions.
[0188] As discussed above, the development of these models becomes application-dependent based on where in the process modeling is attempted. Which parameters are selected to explain chemical heterogeneity can be as simple as only using sequence information.Alternatively, the expression system or process parameters can be introduced with sufficiently large data sets. These models can be used to predict which sequences in conjunction with what set of conditions might minimize intended or unwanted chemical heterogeneity. Alternatively, these models can be used to compare measured mass distributions against what is observed in a kind of a quality control approach. This comparison may be of interest if large data sets are generated under “desirable” conditions, e.g., bioreactor scale or shaker flasks where cells are healthy. A model trained on these data could be compared against molecules (e.g., proteins) generated under high-throughput conditions where cell health or optimal expression conditions may be compromised.What is claimed is:
Claims
CLAIMS1. A computer-implemented method comprising: obtaining mass distributions corresponding to samples produced using a high- throughput expression system; from the mass distributions, determining aligned mass distributions represented on a global mass axis; and causing display of a graphical representation corresponding to the aligned mass distributions.
2. The computer-implemented method of claim 1, wherein determining the aligned mass distributions comprises subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions.
3. The computer-implemented method of claim 1, wherein determining the aligned mass distributions comprises aligning the mass distributions to a single mass distribution of the mass distributions.
4. The computer-implemented method of claim 1, wherein determining the aligned mass distributions comprises minimizing the differences of a mass distribution against all others of the mass distributions.
5. The computer-implemented method of claim 1, wherein determining the aligned mass distributions comprises: determining adjusted mass distributions by subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions; after the subtracting, generating the global mass axis; generating a first matrix including the adjusted mass distributions represented on the global mass axis; identifying, from the first matrix, a correlated group of the adjusted mass distributions; performing an alignment operation of a non-correlated group of adjusted mass distributions against the correlated group; andgenerating a second matrix including the non-correlated group aligned with the correlated group.
6. The computer-implemented method of claim 1, wherein determining the aligned mass distributions comprises: obtaining adjusted mass values by subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions; and performing a modification adjustment on the adjusted mass values.
7. The computer-implemented method of claim 6, wherein the modification adjustment corresponds to a chemical modification of a protein, a polynucleotide, a virus, a virus-like particle, a nanoparticle, a peptide, a lipid, a saccharide, or a metabolite.
8. The computer-implemented method of claim 6, wherein the modification adjustment corresponds to one or more of glycosylation, de-glycosylation, phosphorylation, acetylation, methylation, ubiquitination, sumoylation, lipidation, proteolysis, deamidation, hydroxylation, protein labeling, base insertion, base deletion, or base substitution.
9. The computer-implemented method of claim 1, further comprising generating the global mass axis by finding all unique mass values across the mass distributions, sorting the mass values in increasing order, and using the increase order as the global mass axis.
10. The computer-implemented method of claim 1, further comprising generating the global mass axis by selecting the highest numerical accuracy mass axis from the mass distributions.
11. The computer-implemented method of claim 1, wherein causing display of the graphical representation comprises causing display of one or more of a cluster visualization, a dimensionality reduced visualization, or an underlying factor visualization.
12. The computer-implemented method of claim 1, wherein a first portion of the mass distributions is obtained at a first time point of a sample process and a second portion of the mass distributions is obtained at a second time point of the sample process.
13. The computer-implemented method of claim 12, wherein the first portion and the second portion are obtained from samples of the same substance.
14. The computer-implemented method of claim 1, wherein the samples are biological samples.
15. A system comprising: at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, further cause the system to: obtain mass distributions corresponding to samples produced using a high- throughput expression system; from the mass distributions, determine aligned mass distributions represented on a global mass axis; and cause display of a graphical representation corresponding to the aligned mass distributions.
16. The system of claim 15, wherein the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to subtract the known or expected masses of the samples from the measured masses of the samples in the mass distributions.
17. The system of claim 15, wherein the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to align the mass distributions to a single mass distribution of the mass distributions.
18. The system of claim 15, wherein the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to minimize the differences of a mass distribution against all others of the mass distributions.
19. The system of claim 15, wherein the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to: determine adjusted mass distributions by subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions; after the subtracting, generate the global mass axis; generate a first matrix including the adjusted mass distributions represented on the global mass axis; identify, from the first matrix, a correlated group of the adjusted mass distributions; perform an alignment operation of a non-correlated group of adjusted mass distributions against the correlated group; and generate a second matrix including the non-correlated group aligned with the correlated group.
20. The system of claim 15, wherein the instructions for determining the aligned mass distributions further comprise instructions that, when executed by the at least one processor, further cause the system to: obtain adjusted mass values by subtracting the known or expected masses of the samples from the measured masses of the samples in the mass distributions; and perform a modification adjustment on the adjusted mass values.
21. The system of claim 20, wherein the modification adjustment corresponds to a chemical modification of a protein, a polynucleotide, a virus, a virus-like particle, a nanoparticle, a peptide, a lipid, a saccharide, or a metabolite.
22. The system of claim 20, wherein the modification adjustment corresponds to one or more of glycosylation, de-glycosylation, phosphorylation, acetylation, methylation, ubiquitination, sumoylation, lipidation, proteolysis, deamidation, hydroxylation, protein labeling, base insertion, base deletion, or base substitution.
23. The system of claim 15, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to generate the global mass axis by finding all unique mass values across the mass distributions,sorting the mass values in increasing order, and using the increase order as the global mass axis.
24. The system of claim 15, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to generate the global mass axis by selecting the highest numerical accuracy mass axis from the mass distributions.
25. The system of claim 15, wherein the instructions for causing display of the graphical representation further comprise instructions that, when executed by the at least one processor, further cause the system to cause display of one or more of a cluster visualization, a dimensionality reduced visualization, or an underlying factor visualization.
26. The system of claim 15, wherein a first portion of the mass distributions is obtained at a first time point of a sample process and a second portion of the mass distributions is obtained at a second time point of the sample process.
27. The system of claim 26, wherein the first portion and the second portion are obtained from samples of the same substance.
28. The system of claim 15, wherein the samples are biological samples.
Citation Information
Patent Citations
Method for evaluating data from mass spectrometry, mass spectrometry method, and MALDI-TOF mass spectrometer
US11221338B2
Methods for Mass Spectrometry-Based Structure Determination of Biomacromolecules
US20190018928A1
Alignment of mass spectrometry data
US7365311B1