Method and apparatus for identifying molecular species in a mass spectrum
The method employs regularized regression to optimize mass spectrum coefficients, addressing the challenge of identifying multiple molecular species in chimeric mass spectra, enhancing the accuracy of proteomics analysis.
Patent Information
- Application Number
- JP2023576092
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-07
- Filing Date
- 2022-06-06
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2042-06-06
AI Technical Summary
Existing mass spectrometry methods struggle to accurately identify multiple molecular species in a mass spectrum, particularly in complex samples like unfractionated cell lysates, due to co-isolation and co-fragmentation of precursor ions, leading to chimeric spectra that complicate species identification.
A method and system using regularized regression of mass spectrometry data to optimize mass spectrum coefficients, constraining the number of non-zero coefficients and enforcing sparsity, to identify molecular species in chimeric mass spectra by associating candidate mass spectra with the target spectrum.
Enhances the ability to accurately identify multiple molecular species in mass spectra, improving the precision of proteomics analysis by resolving contributions from individual species even in complex samples.
Smart Images

Figure 0007717849000031 
Figure 0007717849000032 
Figure 0007717849000033
Abstract
Description
Technical Field
[0001] The present invention relates to the identification of molecular species present in a mass spectrum. In particular, it is a method and system for identifying molecular species in a chimeric mass spectrum.
Background Art
[0002] The use of mass spectrometry (MS) technology has become very valuable across many fields where detailed analysis of various chemical samples and often biological samples is required. Such mass spectrometry is used to identify the chemical composition of a given sample.
[0003] Direct analysis in a mass spectrometer typically involves the generation of ions from a chemical sample. These ions are then measured by the mass spectrometer for their mass-to-charge ratio (m / z) and abundance to generate a mass spectrum. The intensity peak (or mass center) at a particular m / z value in such a mass spectrum provides a signature indicating the relative abundance and m / z of each ion. This signature can enable the identification of the compound (or compounds) that make up the original chemical sample.
[0004] Often, for samples containing proteins etc., the identification can be done by automatically searching a library of known mass spectra for known compounds and attempting to match the relative intensity at a given m / z value (within an acceptable range) in the known mass spectrum to the relative intensity at the same m / z value in the measured mass spectrum. If the pattern of relative intensities matches between the two mass spectra, this indicates the presence of the compound corresponding to the library mass spectrum.
[0005] Similarly, a database search approach can be used, where the experimental spectrum is compared to a list of calculated peptide fragment masses derived from a protein database. Such an approach has become an important tool for the identification and quantification of proteins and their modifications in a sample. Such identification is particularly useful with mass spectrometry-based bottom-up proteomics techniques.
[0006] In bottom-up proteomics, typically, proteins are extracted from their origin (e.g., tissue) and can be proteolytically digested using proteases (usually having specific cleavage patterns) to obtain shorter polyamino acid chains called peptides. After digestion, the sample may undergo further processes such as sample washing to remove unwanted salts or other reagents from the sample, offline fractionation (e.g., liquid chromatography) to reduce sample complexity, enrichment of specific classes of (modified) peptides using physicochemical properties or antibodies, or labeling using isobaric tagging reagents (e.g., tandem mass tag TMT, TMTpro, iTRAQ).
[0007] The resulting peptide mixture can then be separated based on their physicochemical properties. For example, a reverse-phase liquid chromatography system directly coupled to a mass spectrometer may be used.
[0008] Peptides eluting from a liquid chromatograph are usually ionized into the gas phase using electrospray ionization (ESI). A mass spectrometer measures the mass-to-charge ratio (m / z) of the peptide ions (or precursor ions). Two operating modes can be distinguished: MS1 and MS2 / MSn-scans. An MS1 or full-MS scan records the m / z of intact precursor ions, and an MS2 / MSn scan isolates a population of ions (e.g., a specific m / z range of an MS1 scan) for subsequent fragmentation (e.g., via collision with neutral gas molecules) and records the m / z of the resulting fragment ions. The specific m / z values and relative intensities of the fragment ions enable the determination of the amino acid sequence. An MS2 / MSn spectrum is often assumed to contain fragments from a single peptide (or precursor). However, depending on the width of the isolation window (the tolerance for selecting ions for a subsequent MS2 / MSn scan), two or more precursor ions can be co-isolated and co-fragmented in one MS2 / MSn event, resulting in a chimeric MS2 / MSn scan containing fragment ions from multiple precursor ions. In fact, in more complex measurement samples such as unfractionated cell lysates on a short gradient, more precursor ions can elute together, and more precursors can elute per unit time such that they are co-isolated in each MS2 / MSn isolation window. This results in an increased likelihood that the resulting MS2 / MSn spectrum contains two or more precursor ions.
[0009] In a mass spectrum generated from a sample containing two or more compounds (such as an MS2 scan having two or more precursor molecular species, or an MS1 scan in which two or more molecular species are present), there can be multiple compounds contributing to the intensity of a given peak. This can make it difficult to identify molecular species from the resulting mass spectrum. Some existing methods aimed at attempting to identify the contribution of individual molecular species to a given mass spectrum include complementary ion boosting as outlined in "Deconvolution of Mixture Spectra and Increased Throughput of Peptide Identification by Utilization of Intensified Complementary Ions Formed in Tandem Mass Spectrometry" (Fedor Kryuchkov et al., Journal of Proteome Research, 12(7), pp. 3362-3371, (2013), DOI: 10.1021 / pr400210m), and DIA-Umpire as outlined in "DIA-Umpire: comprehensive computational framework for data-independent acquisition proteomics" (Tsou, CC. et al., Nature Methods, 12, 258-264 (2015) DOI: 10.1038 / nmeth.3255), etc., but these methods are typically limited to cases where the mass spectrum is from an MS2 scan and is part of a series of mass spectra from a chromatography experiment. Subsequently, the estimated chromatographic features (such as elution peaks) can be used to assist in the identification of precursor ions (typically peptides in a proteomics experiment).
Summary of the Invention
[0010] The object of the present invention is to provide an improved system and method for identifying molecular species in a mass spectrum in which multiple molecular species can be present without the limitations of prior art systems. Such systems and methods can be particularly advantageous for use in the field of proteomics, but it will be understood from the following description that they are not limited to such fields and are generally applicable to mass spectrometry experiments.
[0011] In the present invention, a new method and system for identifying one or more molecular species represented in a mass spectrum are proposed. The present invention provides a method for identifying one or more molecular species represented in a mass spectrum by means of regularized regression of the mass spectrum using a selected candidate mass spectrum corresponding to a candidate molecular species.
[0012] In a first aspect, a method (such as a computer-implemented method) for identifying one or more molecular species represented in mass spectrometry data (such as in the form of a mass spectrum) is provided.
[0013] In this aspect, the method includes obtaining a set of candidate mass spectra (or items of candidate mass spectrometry data) for the mass spectrum (or mass spectrometry data), each candidate mass spectrum (or item of candidate mass spectrometry data) corresponding to a respective candidate molecular species. The method then proceeds to optimize a set of mass spectrum coefficients for the set of candidate mass spectra based on the mass spectrum. The step of optimizing includes varying the mass spectrum coefficient values based on an objective function. The objective function associates (or includes terms associating) a linear combination of the candidate mass spectra with the mass spectrum according to the mass spectrum coefficients. The objective function may include a term that is a function of the difference (or deviation or error) between the linear combination of the candidate mass spectra by the mass spectrum coefficients and the mass spectrum. In some embodiments, the term may be the mean squared error between the linear combination of the candidate mass spectra and the mass spectrum.
[0014] The optimization follows a regularization term of an objective function that constrains the number of non-zero mass spectrum coefficient values. The regularization term may include a norm (such as the L1 norm) of the mass spectrum coefficient values. The regularization term may be configured to enforce sparsity in a vector formed from the coefficient values. In some embodiments, optimizing may further include varying the degree of regularization of the regularization term (e.g., by varying a regularization parameter of the regularization term).
[0015] The method also includes providing, for one or more of the candidate molecular species, respective indicators of matches in a mass spectrum based at least in part on the optimized set of mass spectrum coefficients. Each indicator may include an amount (such as a predicted or estimated amount) of the corresponding candidate present in the mass spectrum.
[0016] The method may further include identifying the mass spectrum as a chimeric mass spectrum based on the optimized set of mass spectrum coefficients.
[0017] In some embodiments, one or more of the molecular species are one or more precursor molecular species, the mass spectrum is a fragment mass spectrum from one or more precursor molecular species, and each candidate mass spectrum is a candidate fragment mass spectrum corresponding to a respective candidate molecular species. The precursor molecular species represented in one or more of the candidate fragment mass spectra may be a peptide.
[0018] The providing step may include identifying, as one or more candidate molecular species, sample molecular species represented in the chimeric mass spectrum based on the optimized set of fragment mass spectrum coefficients.
[0019] In some further embodiments, for a given candidate mass spectrum, a spectral similarity score is calculated between the given candidate mass spectrum and a further mass spectrum. The further mass spectrum is generated by subtracting each of the other candidate mass spectra from the mass spectrum according to an optimized set of mass spectral coefficients.
[0020] In some embodiments, the mass spectrum is an MS1 mass spectrum of one or more molecular species, and each candidate mass spectrum includes an isotope pattern corresponding to a respective candidate molecular species.
[0021] In some embodiments, the one or more molecular species are a plurality of isobarically labeled (such as with tandem mass tags) precursor molecular species, the mass spectrum is a fragment mass spectrum derived from the plurality of isobarically labeled precursor molecular species, and each candidate mass spectrum is a candidate fragment mass spectrum corresponding to a respective candidate isobarically labeled molecular species. The providing step further includes generating reporter ion intensities for one or more of the candidate mass spectra, and generating corrected reporter ion intensities for at least one of the precursor molecular species based on the difference between the reporter ion intensities for the precursor molecular species generated from the mass spectrum and the reporter ion intensities for one or more candidate mass spectra scaled by the corresponding optimized mass spectral coefficients.
[0022] In some embodiments, the precursor molecular species represented in one or more candidate fragment mass spectra are peptides (or peptide precursors such as peptide precursors in a single charged state), the mass spectrometry data is part of a series of mass spectrometry data for separation parameters, and the series of mass spectrometry data is generated from a mixed sample containing a plurality of different isotopically labeled samples. The step of optimizing a set of mass spectral coefficients for a set of candidate mass spectra is repeated for each item of mass spectrometry data in the series of mass spectrometry data to generate, for each item of mass spectrometry data, a respective set of mass spectral coefficients optimized for the value of the separation parameter corresponding to the item of mass spectrometry data. The method further includes obtaining, for each item of mass spectrometry data, a respective set of reporter ion intensities, where each reporter ion intensity in the set corresponds to a respective isotopically labeled sample of the mixed sample, and calculating, for at least one of the peptides (or peptide precursors), the proportion of the abundance of the peptide in the mixed sample corresponding to one of the samples based on the set of reporter ion intensities and the set of optimized mass spectral coefficients.
[0023] In some embodiments, the mass spectrum is an MS1 mass spectrum, and one or more of the candidate mass spectra are selected as part of the set of candidate mass spectra based on the MS2 mass spectrum corresponding to the MS1 mass spectrum.
[0024] In some embodiments, the mass spectrum is an MSn (such as MS2) mass spectrum, and one or more of the candidate mass spectra are selected as part of the set of candidate mass spectra based on another mass spectrum at the same or different mass spectral levels (such as MSn, MSn + 1, or MSn - 1) corresponding to the MSn mass spectrum.
[0025] In some embodiments, the mass spectrum is part of a series of mass spectra for separation (such as chromatographic separation, m / z separation, mass separation, or ion mobility separation), and the regularization term includes constraints that enforce a relationship between the mass spectrum coefficients of a given candidate mass spectrum and the mass spectrum coefficients of the same candidate mass spectrum determined for a series of additional mass spectra (such as additional mass spectra adjacent to the series of mass spectra).
[0026] In some embodiments, the objective function is configured to provide a measure of the difference between a linear combination of candidate mass spectra and a mass spectrum. In some further embodiments, the step of varying is performed for the purpose of obtaining an extreme value of the objective function.
[0027] The present invention also corresponds to and provides an apparatus comprising elements, modules, or components configured to implement the above-described method, such as one or more suitably configured computing devices as described below.
[0028] Accordingly, in particular, the present invention provides a system (or apparatus) for identifying one or more molecular species represented in mass spectrometry data (or in the form of a mass spectrum, etc.). The system includes a candidate selection module configured to obtain a set of candidate mass spectra (or items of candidate mass spectrometry data) for a mass spectrum (or mass spectrometry data), where each candidate mass spectrum (or item of candidate mass spectrometry data) corresponds to a respective candidate molecular species.
[0029] The system further comprises an optimization module configured to optimize a set of mass spectrum coefficients for a set of candidate mass spectra based on the mass spectrum. The optimization includes varying the mass spectrum coefficient values based on an objective function. The objective function associates (or includes terms that associate) a linear combination of the candidate mass spectra with the mass spectrum according to the mass spectrum coefficients. The objective function may include a term that is a function of the difference (or deviation or error) between the linear combination of the candidate mass spectra by the mass spectrum coefficients and the mass spectrum. In some embodiments, the term may be the mean squared error between the linear combination of the candidate mass spectra and the mass spectrum. The optimization follows a regularization term of the objective function that constrains the number of non-zero mass spectrum coefficient values. The regularization term may include a norm (such as the L1 norm) of the mass spectrum coefficient values. The regularization term may be configured to enforce sparsity in a vector formed from the coefficient values. In some embodiments, optimizing may further include varying the degree of regularization of the regularization term (such as by varying a regularization parameter of the regularization term).
[0030] The system also comprises an indicator module configured to provide a respective indication of a match in the mass spectrum for one or more of the candidate molecular species, based at least in part on the optimized set of mass spectrum coefficients. Each indicator may include the amount (such as a prediction or estimate) of the corresponding candidate present in the mass spectrum.
[0031] The present invention also provides one or more computer programs suitable for execution by one or more processors, such computer programs being configured to implement the methods outlined above and described herein. The present invention also provides one or more computer-readable media comprising (or storing thereon) such one or more computer programs, and / or data signals carried via a network. Here, embodiments of the present invention are described by way of example only with reference to the accompanying drawings.
Brief Description of the Drawings
[0032]
Figure 1a
Figure 1b
Figure 1c
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0033] In the following description and drawings, specific embodiments of the present invention are described. However, the present invention is not limited to the described embodiments, and it will be understood that some embodiments may not include all of the features described below. However, it will be apparent that various changes and modifications can be made herein without departing from the broader spirit and scope of the present invention as set forth in the appended claims.
[0034] FIG. 1a schematically shows an exemplary system in which sample 100 is analyzed using a mass spectrometer device 101.
[0035] Sample 100 contains molecules of one or more molecular species. Sample 100 may be obtained directly. In addition to, or instead of, this, sample 100 can be obtained from the output of separation techniques such as liquid (or gas) chromatography, ion mobility separation, etc., as briefly described below. As shown in FIG. 1a, sample 100 is provided to mass spectrometer device 101.
[0036] Mass spectrometer device 101 is configured to receive sample 100. The mass spectrometer device may be configured to ionize sample 100 to generate ions of one or more molecular species. Alternatively, sample 100 may already be ionized, for example as part of the separation technique used to obtain the sample. In some examples, as further described below, mass spectrometer device 101 may be configured to fragment one or more molecular species in sample 100 such that sample 100 contains multiple fragment ions for each molecular species.
[0037] The mass spectrometer device 101 is configured to separate the ions (or fragment ions) of the sample 100 and measure the amount (or relative abundance) of the ions (or fragment ions) present in the sample as a function of the ion (or fragment ion) mass-to-charge (m / z) ratio. The mass spectrometer device 101 may be configured to output the measured values of the ions as mass spectrometry data 190 in the form of a mass spectrum or the like. The operation of such a mass spectrometer device 101 is well known in the art and will not be further described herein. Those skilled in the art will understand that the mass spectrometer device 101 may be of any type. For example, the mass spectrometer may include any one of a time-of-flight (TOF) mass spectrometer, a Fourier transform ion cyclotron resonance mass spectrometer (FT-ICRMS), an Orbitrap (trademark) mass spectrometer, and the like.
[0038] An exemplary graphical representation of the mass spectrometry data 190 in the form of a mass spectrum is also shown in FIG. 1a.
[0039] Mass spectrometry data 190 can be represented in the form of one or more mass channels. Each mass channel corresponds to each mass value (or mass-to-charge ratio, referred to as the m / z value in this specification) at which ion species of the same mass (or mass-to-charge ratio) are detected in the mass spectrometer device 101. Each m / z value corresponds to each ion species and is equal to the molecular mass of each ion species divided by the absolute elemental charge of each ion species. The mass spectrometry data 190 includes one or more intensity values 196-n, and each intensity value 196-n appears for each m / z value (or channel) 194-n. Each intensity value 196-n correlates with the relative abundance (or amount) of the ion species corresponding to each m / z value 194-n as measured by the mass spectrometer device 101 from the sample 100. Each intensity value 196~n may be proportional to the relative abundance of the ion species corresponding to each m / z value. The intensity (or relative abundance) measured by the mass spectrometer device 101 for a given mass channel correlates with the relative abundance of the ion species having the mass (or mass-to-charge ratio m / z) detected by the mass spectrometer device 101. It will be understood that the mass spectrometry data 190 (and thus the mass spectrum) can be represented in any one of several different forms, such as the transient signal generated by an Orbitrap (trademark) mass spectrometer or other Fourier transform mass spectrometer such as an ion cyclotron resonance (FT-ICR) spectrometer, or the ion flight time generated by a time-of-flight (TOF) mass spectrometer. It should be understood that the methods and systems described below can be used with any representation of the mass spectrometry data 190 (or mass spectrum).
[0040] It will be understood that there can be two or more ion species detected in a given mass spectrometry experiment having the same mass (or m / z value). Thus, the intensity value for a given mass channel can represent, as briefly described below, the sum of the abundances of all ion species detected in a given mass spectrometry experiment having the mass (or m / z value) corresponding to that particular mass channel.
[0041] Experimental mass spectrometry data, such as mass spectrometry data 190, may be plotted in the form of a continuum plot indicated by a dashed line and a mass center plot indicated by a solid vertical line. The width of the peak indicated by the dashed line represents the limit of mass resolution, which is the ability to distinguish between two different ion species having similar m / z ratios.
[0042] However, it will be understood that the mass spectrometry data 190 need not be plotted in the form of a graph. In fact, the mass spectrometry data 190 may be represented in any suitable form. For example, the mass spectrometry data 190 may be represented as a list containing one or more intensity values 196-n and one or more m / z values 194-n. In some cases, the mass spectrum may simply be represented as a list of mass centers (or maxima), each mass center being represented as a pair of an m / z value and an intensity value.
[0043] Since there are many techniques commonly used in the art to obtain such mass centers from mass spectrometry data, they will not be further described herein. However, it is understood that the techniques described herein may be performed on a list of mass centers forming the mass spectrometry data or on the raw mass spectrometry data for which suitable techniques are used to identify the intensity maxima (or mass centers).
[0044] FIG. 1b schematically shows a further example in which the sample 100 is analyzed using tandem mass spectrometry. For ease of understanding, FIG. 1b shows a first mass spectrometer device 101 and a second mass spectrometer device 201. However, it should be understood that in tandem mass spectrometry, these typically correspond to a single mass spectrometer device (such as the device described above) operating in a first mode 101 and a second mode 201, respectively. The above description of the mass spectrometer device 101 with respect to FIG. 1a equally applies to these first mass spectrometer device 101 and second mass spectrometer device 201.
[0045] In this further example, as described above in connection with FIG. 1a, an initial sample 100 is provided to a first mass spectrometer device 101 (or a first operating mode). The initial sample 100 contains ions of one or more molecular species, and the ions are separated and measured by the first mass spectrometer device 101 to generate initial mass spectrometry data 190 of the initial sample 100. A second mass spectrometer device 201 (or a second operating mode) is configured to operate such that the mass filter of the mass spectrometer device is switched to a mass selection mode to select a subset of the ions in the initial sample 100 having an m / z ratio within a predetermined range. The subset of ions may be considered to form a further sample 200. This enables ions of different molecular species having similar m / z values to be separated for further analysis as described below. It will be appreciated that the further sample 200 can be formed by selecting ions based on other physicochemical properties such as ion mobility.
[0046] The second mass spectrometer device is configured to fragment the ions in the further sample 200, such that the further sample contains fragment ions 205 for each of the molecular species (known as precursor molecular species) originally present in the further sample 200. The second mass spectrometer device 201 is configured to separate the fragment ions 205 of the further sample 200 and measure the amount of the fragment ions 205 present in the sample as a function of the m / z ratio of the fragment ions 205. In particular, the second mass spectrometer device is configured to generate further mass spectrometry data 290 for the fragmented further sample 200.
[0047] This further example will be understood to be an example of what is generally known as tandem mass spectrometry (or MS / MS). In fact, the second mass spectrometer device 201 is an alternative operating mode of the first mass spectrometer device 101 that is operable to select a further sample 200 from the initial sample 100. Since the precursor ions of the molecular species of the further sample 200 have similar mass-to-charge ratios, the fragmentation of these precursor ions into fragment ions 205 of the further sample 200 enables discrimination between the precursor molecular species based on the expected pattern of the fragment ions 205 associated with each precursor molecular species. This is also advantageous since often the exact mass alone is not sufficient to determine the exact structure of the molecular species. Thus, the second mass spectrometer device 201 may be considered to provide further structural information based on the fragment ions. The mass spectrometric analysis of the first mass spectrometer device 101 (or the first operating mode) is generally known as the MS1 stage, and the mass spectrometric analysis of the second mass spectrometer device 201 (or the second operating mode) is known as the MS2 stage. Further stages (MSn, where n = 3, 4, 5,...) may be added, where the fragment ions 205 from the previous stage (such as the MS2 stage) are again selected based on the m / z ratio and used as precursor ions for the next stage, where they are further fragmented and analyzed in the mass spectrometer device 201, as will be understood. These stages can also be performed by further operating modes of the same mass spectrometer device. This again aids in the discrimination between precursor fragment ions with similar m / z ratios and / or provides further structural information regarding a given precursor ion. In either case, the MS1 mass spectrometer device 101 (or the first operating mode) and the MS2 (or MSn) mass spectrometer device 201 (or the second operating mode) may be considered to be specific examples of the mass spectrometer device 101 described above in relation to FIG. 1a.Examples of such tandem mass spectrometry devices having first and second operating modes include the Orbitrap Q Exactive and Orbitrap Exploris instruments produced by Thermo Fisher Scientific GmbH (Bremen, Germany), both of which include an Orbitrap mass analyzer within the mass spectrometry device. The Orbitrap mass analyzer is described in detail in WO 02 / 078046 (A), the entire content of which is incorporated herein by reference and will not be described in detail here.
[0048] The methods and systems of the present invention described hereinafter in this specification are generally applicable to any of the above mass spectrometry stages and operating modes.
[0049] It will of course be understood that the mass spectrometry devices 101 and 201 may also be used in conjunction with a separation device such as a liquid (or gas) chromatograph. Such a separation device is configured to separate a master sample into a plurality of samples 100. In particular, the separation device is typically configured to elute (or release, or otherwise diverge) the components (or analytes) of the master sample from the separation device as a function of separation parameters (or dimensions). The separation parameter (or parameters) can also be considered an elution parameter, particularly when the separation device includes a chromatograph or chromatography column. For example, the separation device may be a liquid (or gas) chromatograph of a type commonly known in the art. In this example, the elution parameter is the retention time. In other words, it is the duration required for a component to pass through the chromatograph (e.g., the time from when the sample is injected into the device until the component is supplied to the mass spectrometry devices 101 and 201). Since liquid (and gas) chromatographs are well known in the art, they will not be described further herein.
[0050] The analyte is typically released by a separation device as a flow of molecular species and then introduced (or injected) into mass spectrometers 101 and 201, where it may be ionized (although it will be understood that a separation ionization device may be used additionally or alternatively). Thus, sample 100 is typically received by mass spectrometers 101 and 201 as a function of separation parameters. In this way, it will be understood that mass spectrometers 101 and 201 receive samples (or analytes) released simultaneously (or substantially simultaneously) at the same value (or within the same range of values) of elution parameters. As a result, mass spectrometers 101 and 201 may be configured to generate multiple items of mass spectrometry data 190 and 290 (such as in the form of multiple mass spectra) as a function of separation parameters. In other words, each item of mass spectrometry data 190 and 290 (or each mass spectrum) is generated for each respective value of the elution parameter. More specifically, each item of mass spectrometry data 190 and 290 may be considered (or may represent) the mass spectrum of sample 100 released at each respective value of the elution parameter.
[0051] FIG. 1c schematically shows exemplary mass spectrometry data 390 in the form of a mass spectrum that may be generated by either of mass spectrometers 101. 201 (or stage) has been described above in connection with FIGS. 1a and 1b. It will be understood that mass spectrometry data 190 and 290 apply equally to this exemplary mass spectrometry data 390 graphically represented in the form of a mass spectrum, and vice versa. FIG. 1c also schematically shows three further items of mass spectrometry data 390-1, 390-2, and 390-3 graphically displayed in the form of a mass spectrum.
[0052] The exemplary mass spectrometry data 390 shown in FIG. 1c corresponds to samples 100 and 200 containing multiple molecular species. Thus, exemplary mass spectrometry data 390, 201 and 200 from sample 100 generated by mass spectrometer 101 may include contributions from each of the molecular species present in samples 100 and 200.
[0053] Additional mass spectra 390-1, 390-2, 390-3 each correspond to a respective molecular species of the plurality of molecular species in samples 100 and 200. In the example shown in FIG. 1c, there are three molecular species in samples 100 and 200, and thus three additional mass spectra 390-1, 390-2, 390-3 are described. However, it will be understood that any number of molecular species may be present in samples 100 and 200. The additional mass spectra 390-1, 390-2, 390-3 are each mass spectra that correspond to (or represent) a sample containing only the respective molecular species. In other words, the additional mass spectra 390-1, 390-2, 390-3 may each be considered to be (or represent) mass spectra obtained from mass spectrometers 101 and 201 that are derived from samples containing only the respective molecular species. Such additional mass spectra 390-1, 390-2, 390-3 are obtained from mass spectrometers 101 and 201, and although each sample need not be completely pure, it will be understood that the presence of other molecular species in each sample is at a level where the effect on the additional mass spectra can be ignored. In some examples, one or more (or all) of the additional mass spectra 390-1, 390-2, 390-3 may be generated by simulation (or numerical calculation).
[0054] As shown in FIG. 1c, exemplary mass spectrum 390 can be considered to include the sum of further mass spectra 390-1, 390-2, and 390-3. In other words, the exemplary mass spectrum can be (at least partially) composed of the contributions of each of the further mass spectra corresponding to each molecular species in samples 100 and 200. The contribution from each further mass spectrum 390-1, 390-2, 390-3 is scaled by the corresponding scale factor (or coefficient) that represents (or indicates) the proportion of the respective molecular species corresponding to the further mass spectra present in samples 100 and 200. The further mass spectra are typically scaled by multiplying each intensity value in the further mass spectrum by the scale factor for that further mass spectrum.
[0055] As described above, the mass spectrometry data can be considered as a set of intensity values at a specific mass-to-charge ratio (or mass channel). Thus, exemplary mass spectrometry data 390 can be considered to have intensity values at each mass-to-charge ratio (or mass channel) that include the sum of the intensity values of 390-1, 390-2, 390-3 in samples 100 and 200, scaled according to the relative abundances of the molecular species for each further mass spectrum 390-1.
[0056] It will be understood that exemplary mass spectrometry data 390 is obtained from mass spectrometer devices 101 and 201 and that exemplary mass spectrometry data 390 may also include an additional contribution to the intensity values corresponding to measurement error and / or random noise.
[0057] The above mass spectrometry data can be considered in terms of a mathematical function that represents intensity I as a function of mass-to-charge ratio m / z. In this way, exemplary mass spectrum 390, I(m / z), can be given as follows. I(m / z)=c1I1(m / z)+c2I2(m / z)+c3I3(m / z)+η(m / z) where I n, where n = 1, 2, 3 are additional mass spectrometry data 390-1, 390-2, 390-3, and c n is the scale factor for each additional mass spectrometry data, and η(m / z) is a term representing measurement error and / or random noise in mass spectrometry apparatuses 101 and 201, and is used to generate an exemplary mass spectrum 390.
[0058] It will be understood that the above description of the mass spectrometry data and the contribution to the additional mass spectrometry data are applicable to any of the aforementioned mass spectrometry stages. For example, in the case of the MS1 stage or experiment, the exemplary mass spectrometry data 390 includes the intensities of the ions of the molecular species in samples 100 and 200. Each of the additional mass spectrometry data 390-1, 390-2, 390-3 is (or represents) an isotope pattern corresponding to each molecular species in samples 100 and 200. The isotope pattern of a molecular species includes the relative intensities of the ions at each m / z ratio expected from a given sample of that molecular species. In other words, the isotope pattern is understood as a pattern of the relative abundances of several isotopes of a particular molecular species. Therefore, the isotope pattern can enable the calculation of the relative abundance of one isotope of a particular molecular species with respect to another isotope of that particular molecular species. Thus, the isotope pattern of a molecular species can include (or represent) several mass channels, and each mass channel corresponds to a respective isotope of the molecular species. The isotope pattern further includes the respective intensities (or abundances) for each mass channel of the isotope pattern.
[0059] The isotope pattern can be represented (or can include) by a set of intensity scale factors S p for the expected intensity at a given m / z ratio (or mass channel) M p for a given total concentration C of a given molecular species. In other words, the intensity scale factor can follow the relationship M p ∝S p C. When the intensity scale factors are normalized so that they sum to 1, a stronger relationship M p =S pIt will follow C. Another common regularization for the isotope pattern is S j scaling the intensity scale factor such that S j = 1, where M
[0060] is the main or reporter m / z ratio (or mass channel). Depending on the units used for the various quantities and / or the regularization scheme employed, it will be understood that there are numerous mathematically equivalent ways to construct such an intensity scale factor.
[0061] As a further example of an MS2 stage (or a higher MSn stage where n > 2), exemplary mass spectrometry data 390 includes the intensities of fragment ions of precursor molecular species in sample 100. A set of further mass spectrometry data 390-1, 390-2, 390-3 each is (or represents) a fragment mass spectrum corresponding to (or derived from) each respective precursor molecular species in samples 100 and 200. Thus, the further mass spectrometry data 390-1, 390-2, 390-3 in this example each includes the intensity of each predicted fragment ion from the fragmentation of a given precursor molecular species. It will be appreciated that a set of further mass spectrometry data may include the intensities of a subset of the predicted fragment ions from the fragmentation of a given precursor molecular species. As noted above, such a set of further mass spectrometry data can provide the relative abundance of fragment ions for a given precursor molecular species. Thus, the fragment mass spectrum of a precursor molecular species can include (or represent) several mass channels, and each mass channel corresponds to each respective fragment ion generated by the fragmentation of the precursor molecular species. The fragment mass spectrum further includes the respective intensity (or abundance) for each mass channel of the fragment ion mass spectrum.
[0062] From the above description, it is understood that for a given experimental mass spectrum generated by the MS1 stage and a given mass spectrum generated by the MS2 (or higher order) stage, a similar situation can occur where the mass spectrum includes contributions from two or more molecular species (in the case of MS1) that can be referred to as precursor molecular species (in the case of MS2 or higher order). Thus, the resulting mass spectrum can be considered to include a linear combination of the isotope pattern of the molecular species (in the case of MS1) or the fragment mass spectrum (in the case of MS2 and above). In both cases, such a resulting mass spectrum can be considered a chimeric mass spectrum.
[0063] FIG. 2 schematically shows an example of a computer system 1000 that can be used in an embodiment of the present invention. The system 1000 includes a computer 1020. The computer 1020 includes a storage medium 1040, a memory 1060, a processor 1080, an interface 1100, a user output interface 1120, a user input interface 1140, and a network interface 1160, all of which are linked to each other via one or more communication buses 1180.
[0064] The storage medium 1040 can be any form of non-volatile data storage device such as one or more of a hard disk drive, a magnetic disk, an optical disk, a ROM, etc. The storage medium 1040 can store an operating system for the processor 1080 to execute for the computer 1020 to function. The storage medium 1040 can also store one or more computer programs (or software or instructions or code).
[0065] The memory 1060 can be any random access memory (storage unit or volatile storage medium) suitable for storing data and / or computer programs (or software or instructions or code).
[0066] The processor 1080 can be any data processing unit suitable for executing one or more computer programs (such as those stored in the storage medium 1040 and / or the memory 1060). Some of the computer programs, when executed by the processor 1080, can cause the processor 1080 to execute the method according to the embodiments of the present invention and configure the system 1000 as a system according to the embodiments of the present invention. The processor 1080 can include a single data processing unit or multiple data processing units that operate in parallel, separately, or in cooperation with each other. When executing the data processing operations for the embodiments of the present invention, the processor 1080 can store data in the storage medium 1040 and / or the memory 1060 and / or read data from the storage medium and / or the memory. The processor 1080 can include one or more graphics processing units (GPUs) that operate in cooperation with other data processing units of the processor 1080.
[0067] The interface 1100 can be any unit for providing an interface to a device 1220 external to the computer 1020 or removable from the computer. The device 1220 can be one or more of a data storage device, such as an optical disk, a magnetic disk, a solid-state storage device, etc. The device 1220 can have processing capabilities - for example, the device can be a smart card. Thus, the interface 1100 can access data from, provide data to, or interface with the device 1220 according to one or more commands received from the processor 1080.
[0068] The user input interface 1140 is configured to receive input from a user or operator of the system 1000. The user can provide this input via one or more input devices of the system 1000, such as a mouse (or other pointing device) 1260 and / or a keyboard 1240, which are connected to or communicate with the user input interface 1140. However, it will be understood that the user can provide input to the computer 1020 via one or more additional or alternative input devices (such as a touch screen). The computer 1020 can store the input received from the input device via the user input interface 1140 in the memory 1060 for later access and processing by the processor 1080, or can pass it directly to the processor 1080 so that the processor 1080 can respond to the user input accordingly.
[0069] The user output interface 1120 is configured to provide graphical output, visual output, and / or audio output to a user or operator of the system 1000. Thus, the processor 1080 can be configured to form an image / video signal representing the desired graphical output and provide this signal to a monitor (or screen or display unit) 1200 of the system 1000 connected to the user output interface 1120. Additionally or alternatively, the processor 1080 can be configured to form an audio signal representing the desired audio output and instruct the user output interface 1120 to provide this signal to one or more speakers 1210 of the system 1000 connected to the user output interface 1120.
[0070] Finally, the network interface 1160 provides the computer 1020 with the functionality to download data from and / or upload data to one or more data communication networks.
[0071] The architecture of the system 1000 shown in FIG. 2 and described above is merely exemplary, and it will be understood that other computer systems 1000 having different architectures (e.g., having fewer components than shown in FIG. 2, or having additional and / or alternative components than shown in FIG. 2) can be used in embodiments of the present invention. By way of example, the computer system 1000 can include one or more of a personal computer, a server computer, a mobile phone, a tablet, a laptop, a television set, a set-top box, a game console, other mobile devices or household electronic devices, a distributed (or cloud) computing system, and the like.
[0072] FIG. 3 schematically shows the logic apparatus of an exemplary analysis system 400 that can be used to analyze mass spectrometry data 390 such as the mass spectrometry data 190, 290, 390 described above in connection with FIGS. 1a - 1c. The analysis system 400 includes a receiver module 410, a candidate selection module 420, an optimization module 430, and an indicator module 440. The analysis system 400 can be implemented (or embodied) on one or more computer systems such as the exemplary computer system 1000 described above in connection with FIG. 2.
[0073] The receiver module 410 is configured to receive the mass spectrometry data 390. Typically, the receiver module 410 is configured to receive the mass spectrometry data 390 from a mass spectrometer device 101, 201 coupled (or connected) to the analysis system 300. However, it will be understood that the receiver module 410 can be configured to receive the mass spectrometry data 390 from any suitable source including a data storage device, a cloud computing service, a test data generation program, and the like. The mass spectrometry data 390 may be received in the form of a mass spectrum as described above. As described above, the mass spectrometry data 390 has a plurality of mass channels, and each mass channel is filled with (or has) its respective intensity (or intensity value).
[0074] The candidate selection module 420 is configured to obtain a set 490 of candidate mass spectrometry data for the received mass spectrometry data 190. The set 490 of candidate mass spectrometry data may be in the form of a set of candidate mass spectra, as described above. For ease of understanding, candidate mass spectra are used in the following description, but it will be understood that the following considerations are equally applicable to other forms of candidate mass spectrometry data. Each candidate mass spectrum typically corresponds to a respective candidate molecular species. The candidate mass spectra can be obtained from the mass spectrum database 425. In addition to, or instead of, this, it will be understood that some or all of the candidate mass spectra can be generated on demand (or on the fly). In some embodiments, it is understood that the candidate selection module 420 may select all of the candidate mass spectrometry data available in the mass spectrum database 425 as the set of candidate mass spectrometry data.
[0075] A candidate mass spectrum can include a mass spectrum generated by a mass spectrometry experiment for the corresponding molecular species. Preferably, the mass spectrometry experiment has a sample consisting of a pure (or substantially pure) sample of the corresponding molecular species. Additionally, or alternatively, the candidate mass spectrum can include a theoretical (or predicted) mass spectrum. A theoretical (or calculated or predicted) mass spectrum is typically a mass spectrum generated based on a theoretical model of the mass spectrometry experiment. In particular, such a model can be generated using machine learning based on appropriate training data. An example of the generation of a theoretical mass spectrum is the INFERYS system by MSAID GmbH. Other examples include those described in "Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning" (Gessulat, S., Schmidt, T., Zolg, D.P. et al., Nat Methods 16, 509-518 (2019), DOI: 10.1038 / s41592-019-0426-7), "pDeep: Predicting MS / MS Spectra of Peptides with Deep Learning" (Zhou et al., Analytical Chemistry 2017 89(23), 12690-12697 DOI: 10.1021 / acs.analchem.7b02566), and "MS2PIP: a tool for MS / MS peak intensity prediction" (Sven Degroeve, Lennart Martens, Bioinformatics, Volume 29, Issue 24, December 15, 2013, 3199-3203, DOI: 10.1093 / bioinformatics / btt544). Examples of methods for generating theoretical (or calculated) isotope patterns include "Poisson Model To Generate Isotope Distribution for Biomolecules" (Rovshan G. Sadygov, Journal of Proteome Research 2018 17(1), 751-758 DOI: 10.1021 / acs.jproteome.7b00807), "BRAIN: A Universal Tool for High-Throughput Calculations of the Isotopic Distribution for Mass Spectrometry" (Dittwald et al., Analytical Chemistry 2013 85(4), 1991 - 1994 DOI: 10.1021 / ac303439m), "BRAIN 2.0: Time and Memory Complexity Improvements in the Algorithm for Calculating the Isotope Distribution" (Dittwald, P. et al., J. Am. Soc. Mass Spectrom. 25, 588 - 594 (2014) DOI: https: / / doi.org / 10.1007 / s13361-013-0796-5), "Accelerated Isotope Fine Structure Calculation Using Pruned Transition Trees" (Martin Loos et al., Analytical Chemistry 2015 87(11), 5738 - 5744 DOI: 10.1021 / acs.analchem.5b00941), "IsoSpec: Hyperfast Fine Structure Calculator" (Mateusz K. Lacki et al., Analytical Chemistry 2017 89(6), 3272 - 3277 DOI: 10.1021 / acs.analchem 6b01459), "IsoSpec2: Ultrafast Fine Structure Calculator" (Mateusz K. Lacki et al., Analytical Chemistry 2020 92(14), 9472 - 9475 DOI: 10.1021 / acs.analchem.0c00959, and "Fast Exact Computation of the k Most Abundance Isotope Peaks with Layer-Ordered Heaps" (Patrick Kreitzberg et al., Analytical Chemistry 2020 92(15), 10613-10619 DOI:10.1021 / acs.analchem.0c01670). Since the theoretical mass spectrum corresponding to a particular molecular species can be considered pure, it will be understood that the use of the theoretical mass spectrum as a candidate mass spectrum in the invention described herein is particularly advantageous. In other words, the theoretical mass spectrum contains only contributions from the corresponding molecular species. In contrast, experimental mass spectra, such as those obtained from an experimental mass spectrum library, may have contaminated the initial sample used to generate the experimental mass spectrum, or may contain contributions from other molecular species, such as experimental artifacts, noise, etc. As will be apparent from the following description, such contamination in the candidate mass spectrum can increase errors in the subsequent identification of the received mass spectrometry data 390.
[0076] The candidate selection module 420 may be configured to select one or more candidate mass spectra as part of a set 490 of candidate mass spectra based on the received mass spectrometry data 390. For example, the received mass spectrometry data 390 may correspond to (or may be generated using) a given separation window. In other words, samples 100 and 200 may be selected to be used in a mass spectrometer device to generate mass spectrometry data such that only ions within a given mass-to-charge range are analyzed (or detected). In such a scenario, the candidate selection module 420 may be configured to select candidate mass spectra corresponding to candidate molecular species that are within a given separation window (or have a mass-to-charge ratio within the mass-to-charge ratio range). The candidate selection module 420 may further be configured to discard (or not select) candidate mass spectra corresponding to candidate molecular species that are not within a given separation window (or have a mass-to-charge ratio outside the mass-to-charge ratio range). Additionally, or alternatively, the candidate selection module 420 may be configured to select candidate mass spectra based on their similarity to the received mass spectrometry data 390. In particular, the candidate mass spectra may be selected based on a predetermined threshold of peaks within the candidate mass spectra that are present in the received mass spectrometry data 390. For example, the candidate selection module 420 may be configured to discard (or not select) a mass spectrum as a candidate mass spectrum if the mass spectrum has less than a threshold number of peaks within a mass channel that has a peak present in the received mass spectrometry data 390. This can be thought of as selecting candidate mass spectra based on a peak presence score.
[0077] Additionally or alternatively, the candidate selection module 420 may be configured to select candidate mass spectra based on any one or more of the following criteria. · Physicochemical properties of the candidate molecular species A given likelihood that a candidate molecular species is detectable in a mass spectrometer · The separation window of the received mass spectrometry data 390 · The retention time corresponding to the received mass spectrometry data 390 · Charge of the precursor molecular species · Presence and / or absence of peaks in the received mass spectrometry data 390 compared to a given candidate molecular species · Number of shared peaks between the received mass spectrometry data 390 and the candidate molecular species (e.g., the minimum number may be 2 or 3 or more) · Similarity and / or dissimilarity criteria for comparing the intensity of the candidate spectrum with the intensity in the received mass spectrometry data 390 · Regularized spectral contrast angle (i.e., cosine similarity), Pearson correlation coefficient, Spearman correlation, etc.
[0078] The candidate selection module may be configured to select candidate mass spectra based on a mathematical combination (e.g., multiplication, division, etc.) of intensity-based similarity / dissimilarity criteria (e.g., comparing the intensity of the candidate spectrum with the intensity in the received mass spectrometry data 390, cosine similarity, Pearson correlation, Spearman correlation, etc.) and peak presence / absence-based similarity / dissimilarity criteria (e.g., the number of peaks shared between the received mass spectrometry data 390 and the candidate molecular species).
[0079] In addition, or alternatively, the candidate selection module 420 may weight the matches between the peaks present in the mass spectra within the mass channels having peaks present in the received mass spectrometry data 390. Such weighting may be based, for example, on the intensity of the peaks in the received mass spectrometry data 390. In this way, matches to the most abundant peaks in the received mass spectrometry data 390 are given a greater weight. This can be thought of as selecting candidate mass spectra based on a peak intensity score. It is understood that combining intensity-based similarity / dissimilarity criteria with peak presence / absence-based similarity / dissimilarity criteria can mitigate the weaknesses that either class of criteria has when used alone.
[0080] The candidate selection module 420 may also be configured to score candidate mass spectra against the received mass spectrometry data 290 based on a machine learning algorithm, such as linear discriminant analysis (LDA). In particular, the candidate selection module 420 may be configured to select (or discard) a mass spectrum as a candidate mass spectrum based on the LDA score. Such selection may include selecting a mass spectrum having a score exceeding a predetermined threshold. Alternatively, such selection may include ranking the mass spectra by the LDA score and selecting a predetermined number of mass spectra having the highest scores as candidate mass spectra.
[0081] In addition to, or alternatively to, the candidate selection module may be configured to select candidate mass spectra based on information obtained from a further set of the received mass spectrometry data. For example, the received mass spectrometry data 390 may not correspond at all to fragmentation (or may be generated using fragmentation), while a further set of the received mass spectrometry data may correspond to fragmentation (or may be generated using fragmentation). In other words, the mass spectrometry data 390 may correspond to an MS1 spectrum, and a further set of the received mass spectrometry data may correspond to an MS2 spectrum. It will be understood that the two sets of the received mass spectrometry data may correspond to (or be generated using) the same (or similar) separation parameter values. In such a scenario, the candidate selection module 420 may be configured to select candidate mass spectra corresponding to candidate molecular species for the received mass spectrometry data 390 that are within a given window of the separation parameter. Those skilled in the art will understand that classical peak shapes (e.g., Rayleigh distribution, Gaussian distribution) describing the separation of candidate molecular species may be considered for the determination of this window.
[0082] Optimization module 430 is configured to optimize a set of mass spectrum coefficients 432 for a set of candidate mass spectra 490 based on the received mass spectrum. The set of mass spectrum coefficients 432 can be considered to define (or specify) a linear combination of the candidate mass spectra. A mass spectrum coefficient can be assigned (or correspond) to each candidate mass spectrum. Optimization module 430 is typically configured to require (or enforce) that these coefficients are non - negative. The mass spectrum coefficients for a candidate mass spectrum scale the intensity of each peak in the candidate mass spectrum by the value of the mass spectrum coefficient. For example, intensity I p,m Take P candidate mass spectra including a set of, where p = 1, 2,..., P is an index across the mass spectra and m is an index across the mass channels for each candidate mass spectrum. The set of mass spectrum coefficients β p 432 defines a further mass spectrum that is a linear combination of the candidate mass spectra scaled by the corresponding mass spectrum coefficients. This further mass spectrum has a set of intensities.
[0083]
Number
[0084] It will be understood that there may also be a further coefficient β0 that can represent a constant intensity offset for the further mass spectrum. In other words
[0085]
Number
[0086] To optimize a set of mass spectrum coefficients, the optimization module 430 is typically configured to vary each value of the mass spectrum coefficients (or mass spectrum coefficient values) based on an objective function.
[0087] The objective function associates (or includes terms for associating) a linear combination of candidate mass spectra defined by (or according to) a set of mass spectrum coefficients with the received mass spectrum 390. Typically, the objective function provides a measure of the difference between the linear combination of mass spectra and the received mass spectrum 390. The optimization module may be configured to vary each value of the mass spectrum coefficients to obtain an extreme value of the objective function. In particular, when the objective function provides a measure of the difference between the linear combination of mass spectra and the received mass spectrum 390, the values of the mass spectrum coefficients can be changed to minimize the objective function. However, it will be understood that such a minimization problem can be recast as a maximization or as obtaining a saddle point of the objective function by appropriate recasting of the objective function. From the following description, coefficient β p Since a numerical optimization method can be used to obtain the coefficient, it is understood that terms that are substantially negligible with non-linear intensity (such as polynomials) can be included in the objective function without having a perceptible effect on the values of the obtained coefficients. Such an objective function can still be considered as associating a linear combination of candidate mass spectra defined by (or according to) a set of mass spectrum coefficients with the received mass spectrum 390.
[0088] In other words, the optimization module 430 can be considered to be configured to obtain a set of values of the mass spectrum coefficient β p and for each mass channel m,
[0089]
Number
[0090] wherein
[0091] [Number] is the intensity value of the received mass spectrum in mass channel m, and ε m is the error term. The purpose of optimization is to obtain a set of coefficients 432 {β m} such that the set of error terms {ε p} (or a function thereof) is substantially minimized. Thus, it will be understood that the optimization may be a linear regression optimization process. In particular, the sum of the squares of the error terms can be minimized (or sought to be minimized) by optimization.
[0092] Therefore, the objective function can include a term representing the sum of the squares of the differences between the intensity of the received mass spectrum 390 and the respective corresponding intensities of each candidate mass spectrum scaled by the coefficients of the candidate mass spectra. Mathematically, this term can be expressed as follows.
[0093] [Number] where M is the number of mass channels, I m =(I 1,m , I 2,m , I 3,m ,..., I P,m ) is the vector of the intensities of each candidate mass spectrum in the set of candidate mass spectra in mass channel m, and β = (β1, β2, β3,..., β P ) is the vector of mass spectrum coefficients. In some embodiments, it will be understood that the square root of the above term may be used instead.
[0094] It will be understood that other processing of the error terms may be used for minimization instead of the sum of the squares of the error terms described above.
[0095] For example, the mean absolute error can be minimized. Thus, the objective function can include a term representing the sum of the absolute differences between the intensities of the received mass spectrum 390 and the respective corresponding intensities of each candidate mass spectrum scaled by the coefficients of the candidate mass spectra. Mathematically, this term can be expressed as follows.
[0096]
Number
[0097] It will be understood that other processing of the error term can in effect result in the above terms taking the form of one or more existing loss functions. Examples of such loss functions include Huber loss (or smoothed mean absolute error), cosine similarity (or spectral angle), LogCosh loss, maximum likelihood derived loss function (MLE), cross entropy, hinge loss, LINEX loss, and the like.
[0098] The optimization module 430 is also configured to perform such changes in accordance with regularization (or constraints). Regularization is typically enforced by including a regularization (or penalty) term in the objective function. Regularization is configured to constrain (or reduce, or penalize an increase otherwise) the number of non-zero mass spectrum coefficients. The regularization term may be configured to move the objective function away from an extreme value as the number of non-zero coefficients increases. In this way, the regularization term may directly depend on the number of non-zero coefficients. Thus, it will be understood that the regularization term may be configured to enforce sparsity in the vector of coefficients. Additionally or alternatively, the regularization term may indirectly depend on the number of non-zero coefficients. In particular, the regularization term may be proportional to the sum of the magnitudes of the coefficients.
[0099] For example, LASSO (or least absolute shrinkage and selection operator) regularization may be used. In this case, the regularization term can take the following form.
[0100]
Number
[0101] In the case of LASSO regularization, the objective function can take (or include) the following form.
[0102]
Number
[0103]
Number
[0104] It will be understood that in addition to any of the above regularization terms, further regularization terms can be included in the objective function. In particular, an L2 regularization term can be used. An example of such a term is the regularization term used in so-called "ridge regression". When combined with LASSO regularization as described above, this can be called "elastic net" regression. Here, the objective function can take (or include) the following form.
[0105]
Number
[0106] The strength of regularization (such as the value of λ) may be determined in advance, for example, based on previous tests to obtain a desired number of non-zero coefficients. Additionally or alternatively, the optimization module 430 may be configured to perform multiple optimizations while varying the strength of regularization. In this way, the number of non-zero coefficients can be minimized with respect to a measure of the quality of the coefficients. For example, the number of non-zero coefficients can be minimized with respect to the absolute difference in the strength of the linear combination of the mass spectra according to the mass spectral coefficients that do not exceed a predetermined threshold and the received mass spectrum. It will also be understood that other model selection criteria or methods may be additionally or alternatively used to select λ. These examples include any of the Akaike Information Criterion (AIC), Corrected Akaike Information Criterion (AICc), Bayesian Information Criterion (BIC), Adjusted R-squared (R2adj), Cross-Validation (CV), Stepwise Regression, etc.
[0107] It will be understood that the optimization module 430 may be configured to perform changes to the mass spectral coefficients based on the objective function using numerical optimization techniques, many examples of which are known in the art. In particular, the changing may be performed using an iterative method (or procedure) (or may include an iterative method (or procedure), or may be based on an iterative method (or procedure)). Thus, the optimization of the mass spectral coefficients described above may not actually be able to obtain the extreme value of the objective function. The above optimization may be completed (or succeed or end) when a value of the objective function that is appropriately close (or is estimated to be appropriately close) to the extreme value (or the estimated or predicted extreme value) of the objective function is obtained. When the change is performed using an iterative method, the above optimization may be completed when any of the following conditions are met. (a) That a predetermined number of iterations has been exceeded or met (b) That the change in the value of the objective function with respect to the previous iteration is less than a predetermined threshold (c) That the change in the value (or values) of one or more mass spectral coefficients with respect to the previous iteration is less than a predetermined threshold (d) That a predetermined amount of time has elapsed (e) That a predetermined number of processor cycles has elapsed (f) Exceeding (or satisfying) a predetermined maximum number of non-zero coefficients (g) Exceeding (or satisfying) a predetermined maximum ratio of the described deviation (h) The value of the objective function without the regularization term being within a predetermined threshold of its extreme value, etc.
[0108] Typically, the changing is implemented using coordinate descent. However, the changing may be implemented in whole or in part using any of the finite difference methods such as Newton's method, quasi-Newton method, conjugate gradient method, steepest descent method, proximal minimization, subgradient method, proximal gradient method, least angle regression (LARS), quadratic programming, convex programming, etc.
[0109] The optimization module 430 is configured to provide a set of optimized mass spectrum coefficients 432 to the indicator module 440.
[0110] The indicator module 440 is configured to provide respective indications 445 of matches in the received mass spectrum 390 for one or more of the candidate molecular species, based at least in part on the set 432 of optimized mass spectrum coefficients. Typically, the indication 445 is provided to the user, for example, via the user output interface 1120. Additionally or alternatively, the indication 445 may be stored and / or provided to a further system for additional processing.
[0111] It is understood that non-zero mass spectrum coefficients in the set 445 of optimized mass spectrum coefficients may be considered to indicate the presence of the corresponding candidate mass spectrum in the received mass spectrum 390. As a result, non-zero mass spectrum coefficients may be regarded as indicating the presence of the corresponding candidate molecular species in the received mass spectrometry data 390 (or the sample corresponding to the received mass spectrometry data 390). It will also be understood that the magnitude of the mass spectrum coefficients in the set 445 of optimized mass spectrum coefficients may be considered proportional to the abundance (or relative abundance) of the corresponding candidate molecular species in the received mass spectrometry data 390.
[0112] Each instruction may include one or more of a flag (or other marker) that indicates (or signals) the presence of a candidate molecular species within the received mass spectrometry data 390, the relative (or absolute) abundance of the candidate molecular species within the received mass spectrometry data 390, respective optimized mass spectrum coefficients, and the like.
[0113] The indicator module 440 may be configured to indicate the presence of a molecular species when a corresponding mass spectrum coefficient exceeds a predetermined value. In this way, the most likely (or most abundant) molecular species can be indicated (or identified). Additionally or alternatively, the indicator module 440 may be configured to indicate the presence of only a predetermined number of molecular species, for example, those having the largest mass spectrum coefficients (or the largest contribution to the received mass spectrum 390).
[0114] FIG. 4 is a flow diagram schematically showing an exemplary method 480 for identifying one or more molecular species represented in received mass spectrometry data 390 as may be performed by the system 400 of FIG. 3.
[0115] In step 482, which may be performed by the candidate selection module 420, a set 490 of candidate mass spectra for the received mass spectrometry data 390 is obtained. Each candidate mass spectrum corresponds to a respective candidate molecular species. One or more candidate mass spectra are typically selected as part of the set 490 of candidate mass spectra based on the received mass spectrometry data 390. In particular, the candidate mass spectra may be selected based on their similarity (or one or more similarity scores) to the received mass spectrometry data as described above. For received mass spectrometry data 390 generated from an MS2 (or MSn, where n>2) experiment, the received mass spectrometry data 390 may correspond to (or be generated using) a given separation window. In this case, step 482 may include selecting candidate mass spectra of molecular species corresponding to m / z values that fall within the separation window.
[0116] In step 484, a set 432 of mass spectrum coefficients for a set 490 of candidate mass spectra is optimized based on the mass spectrum. Step 484 can be executed by an optimization module 430.
[0117] Step 484 includes step 486. In step 486, the mass spectrum coefficient values are changed based on an objective function. As described above, the objective function associates a linear combination of candidate mass spectra with the received mass spectrometry data 390 according to the mass spectrum coefficients. Step 486 is executed according to a regularization term of the objective function that constrains the number of non-zero mass spectrum coefficient values. As described above, the regularization term may be proportional to the sum of the magnitudes of the coefficients.
[0118] For example, LASSO (or Least Absolute Shrinkage and Selection Operator) regularization may be used. In such an example, the optimizing step 484 can be considered to include solving a minimization problem.
[0119]
Equation
[0120] In step 488, for one or more of the candidate molecular species, an indication of each match in the mass spectrometry data 390 is provided based at least in part on the optimized set 432 of mass spectrum coefficients. Each indication may include one or more of a flag (or other marker) indicating (or signaling) the presence of the candidate molecular species in the received mass spectrometry data 390, the relative (or absolute) abundance of the candidate molecular species in the received mass spectrometry data 390, each optimized mass spectrum coefficient, and the like.
[0121] FIG. 5 is a flowchart schematically showing a further example of an optimizing step such as optimizing step 484 of method 480 shown in FIG. 4.
[0122] In a further example, the regularization term includes a parameter that specifies the degree of regularization, such as the parameter λ described above in connection with the LASSO regularization term.
[0123] In step 510, an initial value of the degree of regularization is obtained. Usually, the initial value is selected as an extreme value determined based on the regularization term used. In particular, the degree of regularization can be selected such that the coefficient is analytically 0. It will be understood that for any change in the coefficient away from 0, if the change in the regularization term is strictly larger than the change in all other terms of the objective function and of the opposite sign, the coefficient can be regarded as analytically 0. In the case of LASSO regularization, this can be achieved when λ is appropriately large. However, it will be understood that sub-optimal initial values may lead to an increase in the computational effort to reach convergence as described below, but any initial value of the degree of regularization may be selected.
[0124] In step 520, the degree of regularization is reduced by a predetermined step value. It will be understood that the size of the step value can be selected based on the desired convergence rate.
[0125] In step 486, the mass spectrum coefficient values are changed based on the objective function using the current degree of regularization as described above. This step 486 of changing results in an optimized set of mass spectrum coefficients being generated for the current degree of regularization.
[0126] Steps 520 and 486 are repeated until the convergence criterion is met (or satisfied). In this way, the degree of regularization changes until convergence of an optimized set of mass spectrum coefficients is obtained (or other stopping criteria are satisfied).
[0127] In step 530, one or more convergence criteria are tested. The convergence (or stop) criteria can include, in addition to, or instead of, exceeding or meeting a predefined number of iterations, the change in the value(s) of one or more mass spectrum coefficients for the previous iteration being less than a predetermined threshold, a predetermined time having elapsed, a predetermined number of processor cycles having elapsed, a minimum number of non-zero coefficients being found, etc.
[0128] It will be appreciated that the iterative process can be accelerated by using, as the starting value for the mass spectrum coefficients in the next step 486 to be varied, the values of the mass spectrum coefficients obtained in the previous iteration.
[0129] FIG. 6 schematically shows an exemplary analysis system 600 according to an embodiment of the present invention. System 600 is the same as system 400 of FIG. 3 except as described below. Accordingly, features common to system 600 and system 400 shall have the same reference numbers and shall not be described again.
[0130] The receiver module 410 is configured to receive a plurality 690 of items of mass spectrometry data 390 (or a plurality of mass spectra). Each received mass spectrum 390 is from a respective sample eluted with the respective value of the separation parameter from a separation device coupled to the mass spectrum device 101, the mass spectrometer devices 101 and 201, the mass spectrum device 201, the mass spectrum device 201 as described above, which is a mass spectrum generated by the mass spectrum device 201. For example, the plurality of mass spectra can correspond to a chromatography experiment having a peptide mixture (such as that obtained by digestion of one or more proteins) as input. In this example, the peptides can elute from the chromatography device as a function of a separation parameter (such as retention time or solvent gradient). Here, each mass spectrum corresponds to a peptide eluted at a particular value of the retention time or solvent gradient.
[0131] Optimization module 430 and candidate selection module 420 are configured to act on each item of mass spectrometry data 390 in substantially the same manner as they act on the mass spectrometry data 690 described in FIG. 4, except for this. In this way, it will be understood that the optimization module can be configured to generate a plurality of sets 632 of optimized mass spectrum coefficients 432, each corresponding to a respective item of mass spectrometry data 390. Similarly, identification module 440 can be configured for a set 645 of indicators 445 corresponding to each item of mass spectrometry data 390.
[0132] In addition, the optimization module may be configured to further constrain the value of the mass spectrum coefficient for the received mass spectrum 390 at one value of the separation value based on the optimized value of the mass spectrum coefficient for the received mass spectrum 390 at a preceding (and / or subsequent) value of the separation parameter. In one example, each coefficient may be penalized based on additional information or criteria.
[0133] In one example, the coefficient β of a given candidate mass spectrum p at one point i within the separation parameter sequence p,i and the coefficient β of the same candidate mass spectrum p at an adjacent (and / or nearby point) point (such as j = i - 1, or j = i + 1, etc.) within the separation parameter sequence p,j may include a penalty term proportional to the difference between them. Here, the index i = 1, 2, 3,... enumerates the sequence of received mass spectra 390 at increasing values of the separation parameter. It will be understood that when combined with the LASSO approach described above, the objective function can take the following form.
[0134]
Equation
[0135]
Number
[0136] FIG. 7 shows an exemplary experimental mass spectrum 790 together with a linear combination of candidate mass spectra calculated according to an embodiment of the present invention.
[0137] The intensity peaks shown below the x-axis on graph 700 represent the experimental mass spectrum 790. The experimental mass spectrum 790 is generated by an MS2 mass spectrometer device supplied with a sample of digested human protein from cells grown in cell culture. The experimental mass spectrum 790 was acquired for a mass of 429.8997 with a tolerance (separation width) of 0.65 Da at a retention time of 48.96 minutes. Higher energy C-trap dissociation (HCD) was used to fragment the molecular species at 28 normalized collision energies. The experimental mass spectrum was obtained using an Orbitrap™ mass analyzer. The spectrum shown in FIG. 6 includes only the experimental peaks (with a specific tolerance) that match the peaks from the predicted spectrum. For clarity of comparison, the intensity peaks corresponding to LYVDFPQHLR a2 ions, phenylalanine immonium ions, and tyrosine immonium ions, which are known to be present in the sample but not considered in the candidate mass spectrum, are omitted. A set of 29 candidate mass spectra including the mass spectra corresponding to the peptides listed in Table 1 was selected, and a set of mass spectral coefficients was generated according to the system and method of the present invention described above.
[0138]
Table 1
[0139] Eleven non-zero coefficients are present within the set of optimized mass spectral coefficients, and four of the largest of these are listed below the heading "Coefficient" in Table 710. The intensity peaks shown above the x-axis correspond to a linear combination of the candidate mass spectra scaled by their respective coefficients for the four largest mass spectral coefficients. The intensity of the combined intensity of the peaks closely resembles the intensity of the matched experimental peaks. Note that the peak 720 of the linear combination includes contributions from multiple candidate mass spectra and exactly reproduces again the corresponding peak in the experimental mass spectrum.
[0140] An additional graph 750 is also shown in FIG. 7. Graph 750 is identical to graph 700 except that the linear combination of the candidate mass spectra includes all 11 candidate mass spectra having non-zero mass spectral coefficients.
[0141] It is understood that the mass spectral coefficients calculated using the above-described system and method can be used in any number of ways to facilitate the analysis of chimeric mass spectra and, further, to identify and quantify the molecular species present in the mass spectrum. For example, the mass spectral coefficients can enable the calculation of an improved spectral similarity score. It will be understood that spectral similarity scores are often used to measure how well two mass spectra match.
[0142] The spectral similarity score is calculated between an experimental spectrum and a theoretical mass spectrum, a predicted mass spectrum, or a previously acquired mass spectrum. An example of a spectral similarity score is a count-based score (or a peak presence / absence-based similarity / dissimilarity score). Such a score may involve determining how many fragment ions match between two spectra. The spectral similarity score may also include a measure of the regularized spectral contrast angle (SA) between the intensities of two spectra. The calculation of spectral similarity scores such as SA is well known to those skilled in the art (see, e.g., Toplak UH et al., "Conserved peptide fragmentation as a benchmarking tool for mass spectrometers and a discriminating feature for targeted proteomics", (Mol & Cell Proteomics. August 2014; 13(8): 2056-71. doi:10.1074 / mcp.O113.036475) and Wan KX et al., "Comparing similar spectra: from similarity index to spectral contrast angle.", (J. Am. Soc. Mass Spectrom. January 2002; 13(1): 85-8. doi:10.1016 / S1044-0305(01)00327-0)), and will not be discussed further herein.
[0143] However, when the experimental spectrum is chimeric, such scores are affected by all the analytes underlying the experimental spectrum, biasing the resulting scores. The calculated mass spectral coefficients described above enable the calculation of an improved spectral similarity score. For example, additional mass spectra can be generated from the experimental mass spectrum by subtracting all contributions of candidate mass spectra, each scaled by the corresponding mass spectral coefficient, except for a given candidate mass spectrum. In other words, using the candidate coefficients, predict all proportional intensities other than a given candidate spectrum, add the proportional intensities together, and then subtract that sum from the experimental spectrum to calculate an additional mass spectrum. This additional mass spectrum can be considered an experimental spectrum from which the contributions of all interfering spectra have been removed. In fact, the additional mass spectrum simulates a situation where the precursor molecular species of a given candidate mass spectrum is the only analyte in the mass spectrum. A spectral similarity score can then be calculated between the additional mass spectrum and the given candidate mass spectrum. In this way, it will be understood that the resulting spectral similarity score is improved because the bias arising from other analytes present can be reduced.
[0144] It is understood that the system described above can also be used for molecular species derivatized with MS2 / MSn instability isotope labeling reagents (e.g., TMT). Molecular species in a given sample 100 can be labeled with a tagging reagent containing a reporter group that generates reporter ions upon fragmentation. Multiple samples labeled with different tagging reagents can then be combined. Identical precursor molecular species that are differentially labeled are separated and fragmented together to obtain an experimental mass spectrum containing one unique reporter ion per sample (or per tagging reagent). The generated reporter ions indicate the relative ratios of the precursor molecular species of the combined samples before pooling the samples and can be used for quantification.
[0145] When the experimental mass spectrum is chimeric, it is understood that such reporter ion intensities can be affected by all molecular species represented in the experimental mass spectrum. Thus, quantification of a given molecular species may be biased due to multiple molecular species in the sample that contribute to the intensity of the same reporter ion. The calculated mass spectral coefficients described above enable the calculation of improved quantitative values for molecular species.
[0146] For example, the optimized mass spectral coefficients can be used to derive chimericity criteria for selecting / deselecting mass spectra to be considered for quantification.
[0147] Additionally or alternatively, the calculated mass spectral coefficients can be used to distribute the experimental reporter ion intensities to the underlying molecular species according to their contributions to the chimeric spectrum by solving a system of linear equations consisting of the corresponding mass spectral coefficients and mixing ratios for different samples obtained from other received mass spectra. In this way, it will be understood that the bias arising from other molecular species present can be reduced, thus improving the quantitative information obtained from the received mass spectrum.
[0148] As an example, in a scenario where there are three experimental spectra, there are two pure spectra generated from peptide A and peptide B respectively, and one chimeric spectrum generated from peptide A, peptide B, and peptide C. To calculate the reporter ion intensity for peptide C in the presence of bias from peptides A and B in the chimeric spectrum, it is possible to multiply the reporter ion intensities of peptides A and B from their corresponding pure spectra by the corresponding optimized coefficients obtained by applying the above method to the chimeric mass spectrum. By subtracting these scaled reporter ion intensities from the reporter ion intensity of the chimeric spectrum, the reporter ion intensity for peptide C is generated with the bias from peptides A and B substantially removed.
[0149] In another example described immediately below, the calculated mass spectral coefficients can be used to determine the fractional abundance of a particular peptide attributable to one of the samples with respect to a mixture of samples.
[0150] In particular, in a separation experiment in which each mass spectrum (or item of mass spectrometry data) is generated for each of a plurality of values of a given separation parameter, the calculated mass spectral coefficients and reporter ion intensities can be obtained for each item of mass spectrometry data. In other words, for each value of the separation parameter (or scan), a set of calculated mass spectral coefficients and a set of reporter ion intensities can be obtained. Since the mass spectral coefficients indicate the abundance of each peptide in the mixed sample for a given scan and the reporter ions indicate the relative proportion of each sample in the mixture present, the proportion of the abundance of the peptides in the mixed sample corresponding to one of the samples can be determined using linear algebra.
[0151] This can be understood in accordance with the explanation outlined below in Attachment A.
[0152] It will be appreciated that the above system can also be used in conjunction with existing target decoy techniques known in the art. Such techniques often involve allowing a number of decoy mass spectra to compete with the fragment mass spectra when attempting to identify fragment mass spectra in experimental mass spectrometry data. Decoy mass spectra are typically appropriately scrambled candidate mass spectra in which the mass channels of the intensity peaks have been altered. Thus, decoy mass spectra do not correspond to "true" precursor molecular species, and any matches between experimental mass spectrometry data and decoy mass spectra must represent false positives (or false discoveries). As such, the number of such decoy matches can be used to estimate the false discovery rate for a given set of experimental mass spectrometry data 390. Thus, this can be used to estimate the false positive likelihood (or other accuracy score) for any matches to non-decoy fragment spectra.General target decoy techniques are well known to those skilled in the art (e.g., Elias, J., Haas, W., Faherty, B. et al., "Comparative evaluation of mass spectrometry platforms used in large-scale proteomics investigations.", (Nat Methods 2, 667-675 (2005). DOI: 10.1038 / nmeth785), Elias JE, Gygi SP. "Target-decoy search strategy for mass spectrometry-based proteomics.", (Methods Mol Biol. 2010, 604:55-71. doi: 10.1007 / 978-1-60761-444-9_5), "Reverse and Random Decoy Methods for False Discovery Rate Estimation in High Mass Accuracy Peptide Spectral Library Searches" (Zheng Zhang, et al. Journal of Proteome Research 2018 17 (2), 846-857, DOI: 10.1021 / acs.jproteome.7b00614).
[0153] Thus, it will be appreciated that the candidate selection modules 420 of systems 400 and 600 can be configured to include one or more decoy mass spectra within a set of candidate mass spectra. The decoy mass spectra can be obtained from a suitable mass spectral database 425. Additionally, or alternatively, the candidate selection module 420 may be configured to generate one or more decoy mass spectra based on one or more existing candidate mass spectra. This may be done using existing known techniques for generating decoy mass spectra. Using theoretical mass spectra as candidate mass spectra, as described above, is particularly advantageous because it can more easily generate decoy mass spectra with realistic intensity distributions.
[0154] It will be understood that the methods described are shown as individual steps to be performed in a particular order. However, one of ordinary skill in the art will understand that these steps can be combined or performed in a different order while still achieving the desired result.
[0155] It will be understood that embodiments of the present invention can be implemented using a variety of different information processing systems. In particular, the figures and their descriptions provide exemplary computing systems and methods, but these are presented only for the purpose of providing a useful reference when explaining various aspects of the present invention. Embodiments of the present invention can be implemented on any suitable data processing device such as a personal computer, laptop, personal digital assistant, mobile phone, set-top box, television, server computer, and the like. Of course, the descriptions of the systems and methods are simplified for purposes of explanation and they are only one of many different types of systems and methods that can be used in embodiments of the present invention. It will be understood that the boundaries between logical blocks are merely exemplary and alternative embodiments may merge logical blocks or elements or impose alternative decompositions of functionality on various logical blocks or elements.
[0156] It will be appreciated that the above functions can be implemented as one or more corresponding modules as hardware and / or software. For example, the above functions can be implemented as one or more software components for execution by a processor of the system. Alternatively, the above functions can be implemented as hardware such as one or more field-programmable gate arrays (FPGAs), and / or one or more application-specific integrated circuits (ASICs), and / or one or more digital signal processors (DSPs), and / or other hardware configurations. Each of the method steps implemented in the flowcharts included herein, or described above, may be implemented by a corresponding respective module. A plurality of the method steps implemented in the flowcharts included herein, or described above, may be implemented together by a single module.
[0157] As long as embodiments of the present invention are implemented by a computer program, it will be understood that a storage medium and a transmission medium carrying the computer program form aspects of the present invention. The computer program can have one or more program instructions, or program codes, which, when executed by a computer, execute embodiments of the present invention. As used herein, the term "program" can be a series of instructions designed to be executed on a computer system, and can include subroutines, functions, procedures, modules, object methods, object implementation forms, executable applications, applets, servlets, source code, object code, shared libraries, dynamic link libraries, and / or other series of instructions designed to be executed on a computer system. The storage medium can be a magnetic disk (such as a hard drive or a floppy disk), an optical disk (such as a CD-ROM, a DVD-ROM, or a BluRay disk), or a memory (such as a ROM, a RAM, an EEPROM, an EPROM, a flash memory, or a portable / removable memory device), etc. The transmission medium can be a communication signal, a data broadcast, a communication link between two or more computers, etc.
[0158] Appendix A As described above, in the homopolymer labeling procedure (such as the tandem mass tagging procedure), a plurality of samples (shown herein as l = L, M, N,...) can each be tagged (or labeled) with a respective tagging reagent containing a respective reporter group that generates a respective reporter ion upon fragmentation.
[0159] Each sample contains one or more peptide species in an unknown proportion. Here, the peptides are shown as a = A, B, C,.... It is understood that the description herein applies to any set of samples containing any set of peptides labeled with an appropriate tagging reagent.
[0160] Typically, the samples are mixed and subjected to several MSn scans, where the scans are denoted by i = 1, 2, 3, ..., Q. In each scan i, the abundance of the reporter ions corresponding to each sample is measured (e.g., as the peak intensity of each reporter ion). The reporter ion intensity for scan i is denoted as R l,i where l is the label of the sample corresponding to the reporter ion. For example, taking the scenario where there are three samples L, M, N in scan i, the intensities R L,i , R M,i , R N,i are obtained.
[0161] Furthermore, the above method is applied to the scan i to calculate (or estimate) the intensity of the peptides present in the mixed sample. In particular, the intensity of any peptide a in a given scan is calculated from the mass spectral coefficient β p corresponding to the candidate mass spectrum of peptide a for that scan, i.e., the index p is equal to a. These calculated intensities of the peptides for scan i are denoted as I a,i where a is the label of the peptide. For example, taking the scenario where there are three peptides A, B, C in scan i, the intensities I A,i , I B,i , I C,i are obtained.
[0162] To complete the notation used herein, it is as follows · The absolute abundance of peptide a in sample l is represented by C l,a · The proportion of the abundance of peptide a for sample l within all samples is
[0163]
Number
[0164] As described in detail below, the above method can be used to determine the intensity of peptides present in the mixed sample in each scan, and the ratio of the peptide abundance for the immobilized peptides between different samples can be calculated. In other words, the ratio of the abundance of each sample to the total abundance of the peptide in all samples can be determined based on the intensity of the peptide present in the mixed sample in each scan.
[0165] Scan i is pure when all I except 1 a,i are 0. For example, pure scan 2 can have I A,2 = 0, I B,2 = 0.5, and I C,2 = 0. Here, scan 2 contains only peptide B, and thus, for all sample 1, R 1,B,2 = R 1,2 and R 1,A,2 = R 1,C,2 = 0. Generally, for a specific peptide a and scan i, when I a,i = 0, then for all sample l, R l,a,i = 0.
[0166] For impure scans, the following equation can be used. In particular, it is understood here that the elution profile is the same for each peptide, regardless of the sample from which it is derived. Thus, the following assumptions can be made. 1. The portions of reporter ions related to a specific peptide and sample sum to the peak intensity of the reporter ions of that sample.
[0167]
Equation
[0168]
Number
[0169]
Number
[0170]
Number
[0171]
Number
[0172]
Number
[0173] From Equation 2, r l,a = R l,a,i / I a,i can be defined, and the following equation is derived.
[0174]
Mathematics
[0175]
Mathematics
[0176] Since the identity matrix I is typically not square, there are at least two approaches to solving the system of linear equations. As is usually the case, if there are more scans (rows) than peptides (columns), the following equation can be set up to obtain r.
[0177]
Mathematics
[0178]
Mathematics
[0179] Alternatively, if there are more peptides than scans, the system of equations can be solved using the constraint
[0180]
Mathematics
[0181]
Mathematics
[0182] Typically, in both cases, the pseudo-inverse of I, i.e., I + can be calculated using known linear algebra techniques, and the solution is given by
[0183] [Mathematics]
[0184] Additionally or alternatively, known numerical techniques such as equation solving using the conjugate gradient method may be used
[0185] [Mathematics] to obtain a solution for.
[0186] For any constant α for a given peptide a a for C l,a = α a r l,a it will be understood that. Thus, the proportion of the abundance of peptide a for sample l within all samples is given by the following equation.
[0187] [Mathematics]
[0188] In this way, the proportion of the abundance that each sample has with respect to the total abundance of the peptide within all samples can be determined. Another aspect of the present invention may be as follows. 〔1〕A method for identifying one or more molecular species represented in mass spectrometry data, obtaining, for the mass spectrometry data, a set of candidate mass spectra, each candidate mass spectrum corresponding to a respective candidate molecular species; optimizing, based on the mass spectrometry data, a set of mass spectral coefficients for the set of candidate mass spectra, the optimizing including varying the mass spectral coefficient values based on an objective function that relates a linear combination of the candidate mass spectra according to the mass spectral coefficients to the mass spectrometry data; adhering to a regularization term of the objective function that constrains the number of non-zero mass spectral coefficient values; and providing, for one or more of the candidate molecular species, respective indications of matches in the mass spectrum based at least in part on the optimized set of mass spectral coefficients. 〔2〕The method according to 〔1〕 above, wherein the one or more molecular species are one or more precursor molecular species, the mass spectrometry data is a fragment mass spectrum derived from the one or more precursor molecular species, and each candidate mass spectrum is a candidate fragment mass spectrum corresponding to a respective candidate molecular species. 〔3〕The method according to 〔2〕 above, wherein the providing step includes identifying the one or more candidate molecular species as sample molecular species represented in a chimeric mass spectrum based on the optimized set of fragment mass spectral coefficients. 〔4〕The method according to 〔2〕 or 〔3〕 above, wherein the precursor molecular species represented in one or more candidate fragment mass spectra are peptides or peptide precursors. 〔5〕The method according to 〔1〕 above, wherein the mass spectrometry data is an MS1 mass spectrum of the one or more molecular species, and each candidate mass spectrum includes an isotope pattern corresponding to a respective candidate molecular species. 〔6〕The method according to 〔1〕 or 〔5〕, wherein the mass spectrometry data is an MS1 mass spectrum, and one or more of the candidate mass spectra are selected as part of the set of candidate mass spectra based on an MS2 mass spectrum corresponding to the MS1 mass spectrum. 〔7〕For a given candidate mass spectrum, 〔2〕The method according to any one of 〔1〕 to 〔6〕, wherein a spectral similarity score is calculated between the given candidate mass spectrum and a further mass spectrum, and the further mass spectrum is generated by subtracting each of the other candidate mass spectra from the mass spectrometry data according to the optimized set of mass spectral coefficients. 〔8〕The method according to any one of 〔1〕 to 〔7〕, wherein each index includes the amount of the corresponding candidate molecular species present in the mass spectrum. 〔9〕The method according to any one of 〔1〕 to 〔8〕, further comprising identifying the mass spectrometry data as a chimeric mass spectrum based on the optimized set of mass spectral coefficients. 〔10〕The one or more molecular species are a plurality of isotopically labeled precursor molecular species, the mass spectrometry data is a fragment mass spectrum derived from the plurality of isotopically labeled precursor molecular species, and each candidate mass spectrum is a candidate fragment mass spectrum corresponding to each candidate isotopically labeled molecular species, 〔6〕The method according to any one of 〔1〕 to 〔4〕, wherein the providing step further includes generating a reporter ion intensity for one or more of the candidate mass spectra, the reporter ion intensity for the precursor molecular species generated from the mass spectrometry data, and the reporter ion intensity for the one or more candidate mass spectra scaled by the corresponding optimized mass spectral coefficient, and generating a corrected reporter ion intensity for at least one of the precursor molecular species based on the difference therebetween. 〔11〕The mass spectrometry data is part of a series of mass spectrometry data for separation parameters, and the series of mass spectrometry data is generated from a mixed sample containing a plurality of samples isotopically labeled in different ways. The step of optimizing a set of mass spectrum coefficients for the set of candidate mass spectra is repeated for each item of mass spectrometry data in the series of mass spectrometry data, and for each item of mass spectrometry data, a respective set of mass spectrum coefficients optimized for the value of the separation parameter corresponding to the item of mass spectrometry data is generated. The method comprises for each item of mass spectrometry data, obtaining a respective set of reporter ion intensities, each reporter ion intensity in the set corresponding to a respective isotopically labeled sample of the mixed sample. The method according to [4], further comprising calculating, for at least one of the peptide precursors, a proportion of the abundance of the peptide in the mixed sample corresponding to one of the samples based on the set of reporter ion intensities and the set of optimized mass spectrum coefficients. 〔12〕The method according to any one of [1] to
[11] , wherein the mass spectrometry data is part of a series of mass spectrometry data for a separation parameter, and the regularization term includes a constraint that enforces a relationship between the mass spectrum coefficients for a given candidate mass spectrum and the mass spectrum coefficients of the same candidate mass spectrum determined for the series of further mass spectrometry data. 〔13〕The regularization term includes the L 1 norm of the mass spectrum coefficient values, and optionally, the regularization term includes the L 2 norm of the mass spectrum coefficient values. The method according to any one of [1] to
[12] . 〔14〕The method according to any one of [1] to
[13] , wherein the optimizing step further includes varying a parameter that specifies a degree of regularization of the regularization term. 〔15〕The objective function provides a measure of the difference between the linear combination of the candidate mass spectra and the mass spectrometry data, and optionally, the varying step is performed for the purpose of obtaining an extreme value of the objective function. The method according to any one of [1] to
[14] . 〔16〕A system configured to perform the method according to any one of [1] to
[15] . 〔17〕A computer program that, when executed by a processor, causes the processor to perform the method according to any one of [1] to
[15] . A computer-readable medium storing the computer program according to
[17] .
Claims
1. A method for identifying one or more molecular species represented in mass spectrometry data, comprising: obtaining, for the mass spectrometry data, a set of candidate mass spectra, each candidate mass spectrum corresponding to a respective candidate molecular species; optimizing, based on the mass spectrometry data, a set of mass spectral coefficients for the set of candidate mass spectra, the optimizing comprising: changing values of the mass spectral coefficients based on an objective function that associates a linear combination of the candidate mass spectra with the mass spectrometry data according to the mass spectral coefficients; adhering to a regularization term of the objective function that constrains the number of non-zero mass spectral coefficients; providing, for one or more of the candidate molecular species, respective indications of matches in the mass spectrometry data based at least in part on the optimized set of mass spectral coefficients.
2. The method of claim 1, wherein the one or more molecular species are one or more precursor molecular species, the mass spectrometry data is a fragment mass spectrum derived from the one or more precursor molecular species, and each candidate mass spectrum is a candidate fragment mass spectrum corresponding to a respective candidate molecular species.
3. The method of claim 2, wherein the providing step comprises identifying the one or more candidate molecular species as sample molecular species represented in a chimeric mass spectrum based on the optimized set of fragment mass spectral coefficients.
4. The method of claim 2, wherein the precursor molecular species represented in the one or more candidate fragment mass spectra are peptides or peptide precursors.
5. The method of claim 1, wherein the mass spectrometry data is an MS1 mass spectrum of the one or more molecular species, and each candidate mass spectrum includes an isotope pattern corresponding to a respective candidate molecular species.
6. The method according to any one of claims 1 to 5, wherein the mass spectrometry data is an MS1 mass spectrum, and one or more of the candidate mass spectra are selected as part of the set of candidate mass spectra based on an MS2 mass spectrum corresponding to the MS1 mass spectrum.
7. For a given candidate mass spectrum, A spectral similarity score is calculated between the given candidate mass spectrum and a further mass spectrum, the further mass spectrum being generated by subtracting each of the other candidate mass spectra from the mass spectrometry data according to the optimized set of mass spectral coefficients, the method according to any one of claims 1 to 5.
8. The method according to any one of claims 1 to 5, wherein each indicator comprises the amount of the corresponding candidate molecular species present in the mass spectrum.
9. The method according to any one of claims 1 to 5, further comprising identifying the mass spectrometry data as a chimeric mass spectrum based on the optimized set of mass spectral coefficients.
10. The one or more molecular species are a plurality of isotopically labeled precursor molecular species, the mass spectrometry data is a fragment mass spectrum derived from the plurality of isotopically labeled precursor molecular species, and each candidate mass spectrum is a candidate fragment mass spectrum corresponding to each candidate isotopically labeled molecular species. The providing step further comprises generating a reporter ion intensity for one or more of the candidate mass spectra, the reporter ion intensity for the isotopically labeled precursor molecular species generated from the mass spectrometry data, and the reporter ion intensity for the one or more candidate mass spectra scaled by the corresponding optimized mass spectral coefficients, and generating a corrected reporter ion intensity for at least one of the isotopically labeled precursor molecular species based on the difference between the reporter ion intensities. The method according to any one of claims 1 to 4.
11. The mass spectrometry data is part of a series of mass spectrometry data for separation parameters, the series of mass spectrometry data being generated from a mixed sample containing a plurality of samples isotopically labeled in different ways. Optimizing a set of mass spectral coefficients for the set of candidate mass spectra is repeated for each item of mass spectrometry data in the series of mass spectrometry data to generate a respective set of optimized mass spectral coefficients for the value of the separation parameter corresponding to the item of mass spectrometry data for each item of mass spectrometry data. The method comprises For each item of the mass spectrometry data, obtaining each set of reporter ion intensities, wherein each reporter ion intensity in the set corresponds to each isotopically labeled sample of the mixed sample; The method according to claim 4, further comprising calculating, for at least one of the peptide precursors, a ratio of the abundance of the peptide in the mixed sample corresponding to one of the samples based on the set of reporter ion intensities and the set of optimized mass spectrometry coefficients.
12. The method according to any one of claims 1 to 5, wherein the mass spectrometry data is part of a series of mass spectrometry data for separation parameters, and the regularization term includes a constraint that enforces a relationship between the mass spectrometry coefficients for a given candidate mass spectrum and the mass spectrometry coefficients of the same candidate mass spectrum determined for the series of additional mass spectrometry data.
13. The regularization term is the L 1 norm of the mass spectrum coefficients, the method according to any one of claims 1 to 5.
14. The method according to claim 13, wherein the regularization term includes the L 2 norm of the mass spectrum coefficient. 2 norm of the mass spectrum coefficient.
15. The method according to any one of claims 1 to 5, wherein the optimizing further comprises varying a parameter that specifies a degree of regularization of the regularization term.
16. The method according to any one of claims 1 to 5, wherein the objective function provides a measure of the difference between the linear combination of the candidate mass spectra and the mass spectrometry data.
17. The method according to claim 16, wherein the varying is performed for the purpose of obtaining an extreme value of the objective function.
18. A system for identifying one or more molecular species represented in mass spectrometry data, the system comprising a candidate selection module configured to obtain a set of candidate mass spectra, each candidate mass spectrum corresponding to a respective candidate molecular species; An optimization module configured to optimize a set of mass spectrometry coefficients for the set of candidate mass spectra based on a mass spectrum, the optimization including varying values of the mass spectrometry coefficients based on an objective function, the objective function associating a linear combination of candidate mass spectra with the mass spectrometry data according to the mass spectrometry coefficients and following a regularization term of the objective function that constrains the number of non-zero mass spectrometry coefficients; An index module configured to provide each indication of a match in mass spectrometry data for one or more of the candidate molecular species; A system comprising the same.
19. A computer program that, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 5.
20. A computer-readable medium storing a computer program that, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Methods for generating, searching and statistically validating a peptide fragment ion library
EP3002696A1
Prediction device, prediction method, and prediction program
JP2016139336A
Substance identification method using mass analysis and mass analysis data processing device
JP2018119897A
Impact surface for improved ionization
JP2018508964A
Methods for Generating Local Mass Spectral Libraries for Interpreting Multiplexed Mass Spectra
US20140138537A1