Physico-chemical property scoring for structure identification in ion spectroscopy

By providing matching scores of mobility or related properties for candidate molecular structures, the problem of difficulty in associating molecular structures with spectral data signal peaks in the existing technology is solved, and the accuracy and reliability of protein identification are improved.

CN115436347BActive Publication Date: 2025-10-17BRUKER SCIENTIFIC LLC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202111042605.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-09
Filing Date
2021-09-07
Publication Date
2025-10-17
Estimated Expiration
2041-09-07

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively associating molecular structures with signal peaks in spectral data obtained from separation based on physicochemical properties, especially in the case of spectral ambiguity, making protein identification difficult.

Method used

By calculating, estimating or inferring a matching score that provides mobility or related properties for each candidate molecular structure, combining the experimental values ​​of multiple physicochemical properties, using the assumed matching scores and possible distributions to distinguish between credible and implausible associations, the matching process between molecular structures and signal peaks is optimized.

Benefits of technology

The accuracy of the correlation between molecular structure and spectral data signal peaks is improved, the reliability and accuracy of protein identification are enhanced, and the impact of spectral ambiguity is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115436347B_ABST
    Figure CN115436347B_ABST
Patent Text Reader

Abstract

The invention relates to a method of associating a molecular structure with a signal peak in spectral data obtained from separation according to one or more physico-chemical properties, the method comprising, as appropriate and possibly iteratively: providing one or more signal peaks in the obtained spectral data that are associated with experimental values of a mobility or a related property; determining one or more candidate molecular structures suitable for association with the one or more signal peaks; providing, for each candidate molecular structure, a first distribution of matching scores as a function of mobility using one of calculation, estimation, derivation and extrapolation; defining, for each candidate molecular structure, a first assumed matching score as output from the corresponding distribution upon application of the experimental values of the mobility or the related property of the one or more signal peaks; and using the first assumed matching score in the step of associating a molecular structure with the one or more signal peaks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to a method of associating a molecular structure with a signal peak contained in spectral data obtained from a separation according to one or more physico-chemical properties. The spectral data can comprise ion spectral data. The ion spectral data can be obtained from a separation according to a physico-chemical property in the gas phase, such as the ion mobility K or the ion mass m or the ion mass-to-charge ratio m / z. The ion spectral data can also be obtained from a separation according to a physico-chemical property before ionization and / or transfer to the gas phase, such as the retention time in liquid chromatography. BACKGROUND

[0002] The relevant art is described in this specification in connection with specific aspects. However, it should not be understood that a description of the relevant art in this respect limits the disclosure below of the present invention. Useful developments known in the relevant art and modifications thereof can also be applied beyond the relatively narrow scope of this specification, and will be apparent to those skilled in the art after reading the disclosure of the present invention as set forth below.

[0003] Database search algorithms for identifying proteins have advanced since the first introduction of the SEQUEST software package in 1994 (Jimmy K. Eng et al., J. Am. Soc. Mass Spectrom. 1994, 5, 976-989). Search engines compute scores by comparing experimental spectra from MS / MS data to in-silico spectra from a protein database. Despite the significant efforts invested in improving protein identification, there is still a large number of spectra that cannot be identified. One of the obstacles is the ambiguity of the spectra. For example, multiple candidate peptides from a protein database can often have the same precursor mass and similar fragment ion patterns, resulting in competitive peptide spectrum match search scores.

[0004] In the following, a brief description of publications dealing with the evaluation of such existing types of spectra will be given, but without requiring completeness:

[0005] For example, the patent document WO 2004 / 013635 A2 proposes a system and method for scoring peptide matches. Embodiments include introducing a scoring system based on proper signal detection, including generating a stochastic model based on one or more matching features related to a first peptide, a second peptide, and fragments thereof, to score a match between the first peptide and the second peptide. A first probability of a match between the first peptide and the second peptide is calculated based on the stochastic model, and a second probability of a non-match between the first peptide and the second peptide. And the match between the first peptide and the second peptide is scored based at least in part on a ratio between the first probability and the second probability.

[0006] By way of further example, patent document WO 2020 / 016428 Al discloses a method for determining the identity of at least one entity from a mass spectrum of the entity and optionally from additional data from chemical, physical, biochemical or biological analysis of the entity, comprising for each entity the steps of: a) collecting analytical data from the mass spectrum of the entity and optionally additional analytical data from chemical, physical, biochemical or biological analysis of the entity; b) obtaining a plurality of candidate identities for the entity and obtaining a prevalence of the candidate identities for the entity, whereas for each candidate identity all candidate identities having a higher prevalence are included in the plurality of candidate identities; c) for each candidate identity of the entity, calculating a score thereof, the calculation involving at least the prevalence of the entity, or at least the prevalence of the entity and the conformity with the mass spectrum, d) determining the entity identity as the candidate identity having a score closest to the score corresponding to the true identity of the entity.

[0007] In the further course of scientific development, solutions for the precise resolution of large numbers of spectra have emerged by adding an additional dimension of separation in the measurement cycle, which dimension has hitherto usually included the retention time from liquid chromatography and the ion mass or mass-to-charge ratio m / z from (tandem) mass spectrometry. The additional dimension of separation is the CCS (collision cross section) originating from an ion mobility separation stage, optionally connected with upstream and / or downstream mass separation stages and / or upstream chromatographic separation stages.

[0008] In the following, a brief description of publications dealing with the evaluation of such enhanced spectra will be given, again without claiming completeness:

[0009] Sheila C. Henderson et al. (Anal. Chem. 1999, 71, 291-301) report a method for the identification of a number of different peptide sequences with indistinguishable molecular weights by an integrated approach, including the ability to resolve different oligomer sequences by their differences in mobility and the assignment of peaks based on comparison with molecular models.

[0010] Philip D. Mosier et al. (Anal. Chem. 2002, 74, 1360-1370) report quantitative structure-property relationships (QSPR) to predict ion mobility spectrometry (IMS) collision cross sections for singly protonated lysine-terminated peptides using information from topological molecular structure and various amino acid parameters. It was found that a first-order amino acid sequence alone was sufficient to accurately predict the collision cross section. These models were constructed using multiple linear regression (MLR) and computational neural networks (CNN).

[0011] Patent documents WO 2010 / 119289 A2 and WO 2010 / 119293 A2 respectively disclose a method of estimating molecular cross-sectional area for predicting ion mobility, giving gas phase interaction radius determination and cross-sectional algorithm calculations to provide separation and characterization of structurally related isomers.

[0012] Patent document WO 2011 / 027131 A1 discloses a method of screening a sample to determine whether one or more known compounds of interest are present. A fragmentation device repeatedly switches between a fragmentation mode of operation and a non-fragmentation mode of operation. It is determined whether a candidate target parent ion (M1) is present in the non-fragmentation dataset and whether one or more corresponding target fragment ions (M2, M3, M4...) are present in the fragmentation dataset. It is further determined whether the candidate target parent ion (M1) and the one or more corresponding target fragment ions (M2, M3, M4...) have substantially similar elution or retention times and / or ion mobility drift times.

[0013] Stephen J. Valentine et al. (J Proteome Res. 2011 May 6; 10(5) 2318-2329) report a scoring scheme using drift times obtained from a drift tube gas phase ion mobility separator to aid in the identification of peptide ions.

[0014] Patent document WO 2011 / 128703 A1 discloses a method and apparatus for identifying and / or characterizing a sample that can incorporate two or more isomeric or isobaric compounds, such as hydroxylated metabolites. The method includes modeling a set of possible structures for each of two or more known isomers or isobaric compounds, calculating a theoretical collision cross section for each modeled structure, and averaging the calculated values for each known compound to provide a value for the theoretical collision cross section for each known compound. A traveling wave ion mobility cell is used to measure a collision cross section value for a compound of the sample, which is then compared to the theoretical value to determine which of the two or more known compounds the compound of the sample most closely resembles.

[0015] Patent document WO 2014 / 170664 A2 discloses a method of screening a sample for at least one compound of interest. The method includes comparing an ion mobility and at least one additional physicochemical property of an ion of the target compound to the same properties of candidate ions in the sample. The properties of the target compound match the properties of a candidate ion in the sample, and it is then determined that the sample contains the target compound.

[0016] Patent document WO 2015 / 136272 A1 discloses a mass spectrometric analysis method in which one or more collision or interaction cross sections of analyte ions are experimentally determined or measured and the mass or mass-to-charge ratio, and a first list of possible candidate compounds is compiled from the determined or measured mass or mass-to-charge ratio. A collision or interaction cross section is then theoretically calculated, estimated or determined for each candidate compound in the first list. The theoretically calculated, estimated or determined collision or interaction cross section is then used to filter or remove candidate compounds from the first list or to reduce a likelihood value associated with one or more candidate compounds in the first list.

[0017] In view of the above, there remains a need for an improved method of associating molecular structures with signal peaks in spectral data obtained from a separation according to one or more physico-chemical properties. Further advantages and benefits of the technical teaching will be apparent to the skilled person upon reading the following disclosure. SUMMARY

[0018] In a first aspect, the invention relates to a method of associating molecular structures with signal peaks in spectral data obtained from a separation according to one or more physico-chemical properties, comprising, as appropriate and repeatable:

[0019] providing one or more signal peaks in the acquired spectral data that are related to an experimental value of (gas phase ion) mobility or a related property;

[0020] determining one or more candidate molecular structures that are suitable for association with the one or more signal peaks;

[0021] providing for each candidate molecular structure a (individual) distribution of first match scores as a function of mobility or related property by one of calculating, estimating, deriving and inferring;

[0022] defining for each candidate molecular structure a putative first match score as an output from the respective distribution upon application of the experimental value of mobility or related property of the one or more signal peaks; and

[0023] using the putative first match score in the step of associating molecular structures with the one or more signal peaks.

[0024] In various embodiments, the putative first match score can be used to exclude molecular structures from the association. A landmark can be established on the scale of match scores for the first match score, which defines a first range that indicates exclusion of molecular structures from the association and a second range that indicates a potential true association.

[0025] In various embodiments, one or more signal peaks can have one or more experimental values of a second physico-chemical property and each candidate molecular structure can be associated with one or more candidate values of the second physico-chemical property. The one or more candidate values can show a level of agreement with the one or more experimental values of the second physico-chemical property, thereby implying a second matching score for each candidate molecular structure. The second matching score can also be used in the step of associating molecular structures with one or more signal peaks. The second matching score and the assumed first matching score can be used separately but jointly in a (quadratic) discriminant analysis in order to discriminate between likely matches or identifications and unlikely matches or identifications.

[0026] In various embodiments, the assumed first matching score and the second matching score can be combined to generate a third matching score. The third matching score can be used in the step of associating molecular structures with one or more signal peaks. Preferably, the candidate molecular structure having the most extreme value of at least one of the assumed first matching score, the second matching score and the third matching score can be used to associate molecular structures with one or more signal peaks. In particular, the candidate molecular structure having the highest score of at least one of the assumed first matching score, the second matching score and the third matching score can be used to associate molecular structures with one or more signal peaks. Furthermore, all candidate molecular structures having a lower score relative to the highest score described above can be considered as not indicating a valid association.

[0027] In various embodiments, the one or more experimental values of the second physico-chemical property and the one or more candidate values of the second physico-chemical property associated with each candidate molecular structure can indicate the molecular mass of at least one of a precursor ion species and a related fragment ion species of the precursor ion species upon dissociation. The fragment ion species can be generated by dissociating a single precursor ion species, or a plurality of fragment ion species can be generated by simultaneously dissociating two or more precursor ion species. In the latter case, the number of precursor ion species dissociated together can be controlled by the settings of the mobility separation stage and additional mass filtering occurring between the mobility separation stage and the dissociation or fragmentation stage.

[0028] In various embodiments, the separation according to the (gas phase ion) mobility or related property is at least preceded or followed by a separation according to a second physical-chemical property. Preferably, the separation according to the second physical-chemical property can comprise at least one of mass or mass-to-charge ratio (m / z) filtering and mass or mass-to-charge ratio dispersion, e.g. time-of-flight (TOF) dispersion in a flight tube. The flight tube can be substantially field-free or can contain field-free regions and reflectrons. Mass or mass-to-charge ratio filtering can in particular provide information that can be used to determine a subset of potential candidate molecular structures from a large number of generally available candidate molecular structures that would otherwise be unmanageable due to their sheer size, in particular for specific classes of molecules of interest, such as peptides and proteins, etc., where millions of molecular structures are conceivable.

[0029] In various embodiments, each distribution can be configured such that it can result in a first match score that deviates from each other. Preferably, the first match score that can be produced by applying (or inserting) an experimental or experimentally determined mobility or related property value to (or into) the distribution can deviate as a function of the distance from a mobility or related property seed value of the distribution, e.g. characterizing the center or the position of highest score of the distribution. Each distribution can be separate, e.g. each distribution can be characterized by separate and explicit parameters, e.g. a seed value locating the center or the position of highest score and a width value indicating the rate of decline as a function of the distance from the center or the position of highest score.

[0030] In various embodiments, each distribution can be configured such that it can lead to a region with the highest first match score and a (first) neighboring region with a reduced first match score adjacent to it along the mobility or related property scale. Preferably, each distribution can also be configured such that it can lead to a second neighboring region with a reduced first match score relative to the region with the highest first match score and the first neighboring region with a reduced first match score. In particular, the distribution can exhibit three or more distinct first match scores depending on the positioning of the experimental values of the mobility or related property. Preferably, each distribution can follow an analytical function, such as a Gaussian function or other suitable mathematical function. For example, such mathematical function can be defined by one or more mobility or related property seed values and one or more width parameters. In the case of a Gaussian function, the width parameter can be a sigma (σ) width or a full width at half maximum (FWHM) width. The sigma width can be twenty percent (20), fifteen percent (15), ten percent (10), five percent (5), or one percent (1) of the mobility or related property seed value (e.g., the center value or the highest score value). The function can be substantially continuous or can be stepwise continuous. Typically, the distribution can be configured to exhibit, for example, a single apex. In addition, the distribution can be configured to be symmetric such that it has the same form or shape in either direction along the mobility or related property scale starting from the mobility or related property seed value, e.g., the center value of a Gaussian function. The distribution can likewise be configured to be asymmetric such that it has different forms or shapes depending on the direction along the mobility or related property scale starting from the seed value. In other words, the distribution can also be skewed or distorted. The first distribution of the first candidate molecular structure and the second distribution of the second candidate molecular structure can partially overlap.

[0031] In various embodiments, the first match score can indicate a probability on a scale between a first value (exclusion match) and a second value (match determination). More generally, the first match score can be a scalar. In particular, the first value can equal zero and the second value can equal one.

[0032] In various embodiments, the calculating, estimating, deriving, or inferring can comprise a method based on at least one of (i) statistical evaluation, (ii) machine learning, and (iii) deep learning of previously acquired and characterized spectral measurement datasets. The distribution of the first match score can represent a standard deviation (or other parameter indicative of reproducibility) among a plurality of existing spectral datasets of a particular candidate molecular structure. If no previous spectral data of a particular candidate molecular structure exists, the distribution can likewise represent an estimation, regression, interpolation, or extrapolation. In various embodiments, the machine learning or deep learning can be performed using a mixture density network (MDN) model.

[0033] In various embodiments, one or more signal peaks can be produced by an ion species of a biomolecule origin. Preferably, the ion peaks can be produced by a peptide, a protein, a lipid, a glycan, a polysaccharide, an oligonucleotide, a metabolite, etc.

[0034] In various embodiments, one or more candidate molecular structures can be determined from a pool of target candidates indicative of possible molecular structures and a pool of decoy candidates indicative of impossible molecular structures. A putative first match score can be used to define a metric that helps to distinguish between a trustworthy association and an untrustworthy association, especially with the knowledge that a decoy candidate match is impossible to be true.

[0035] In a second aspect, the present invention relates to a method of associating a molecular structure with a signal peak in spectral data obtained from a separation performed according to one or more physico-chemical properties, comprising:

[0036] providing a plurality of signal peak groups and a plurality of experimental values of (gas phase ion) mobility or related properties in the acquired spectral data, each signal peak group being associated with an experimental value of (gas phase ion) mobility or related property and having one or more signal peak values;

[0037] determining a plurality of candidate molecular structure groups from a pool of target candidates indicative of possible molecular structures and a pool of decoy candidates indicative of impossible molecular structures, each candidate molecular structure group having one or more candidate molecular structures and being suitable for association with one or more signal peak groups;

[0038] providing one or more candidate values of (gas phase ion) mobility or related properties for each candidate molecular structure by calculating, estimating, inferring and deducing;

[0039] providing a plurality of match scores by defining, for each candidate molecular structure, one or more match scores as a function of a level of agreement between the one or more candidate values of (gas phase ion) mobility or related properties and the experimental values of (gas phase ion) mobility or related properties of the plurality of signal peak groups, and

[0040] using the plurality of match scores to define a metric that helps to distinguish between a trustworthy association and an untrustworthy association.

[0041] In various embodiments, the match score can be a scalar. Preferably, the scalar can take values between zero and unity.

[0042] In various embodiments, a match score scale can be established with match score landmarks that define: a first range that is assumed to be indicative of an untrustworthy association, regardless of whether the underlying candidate molecular structure is from the pool of decoy candidates or the pool of target candidates; and a second range that is assumed to be indicative of a trustworthy association.

[0043] Preferably, the match score threshold can be defined as the percentage of the signal peaks associated with the candidate molecular structure in the decoy candidate pool that is less than one of five percent (5), four percent (4), three percent (3), two percent (2), and one percent (1) of the signal peaks associated with the candidate molecular structure in the decoy candidate pool is located in the second range.

[0044] In a third aspect, the present application relates to a device for recording ion species separated according to one or more physico-chemical properties, comprising a data processing unit designed and configured for carrying out a method as described above. BRIEF DESCRIPTION OF DRAWINGS

[0045] The application can be better understood with reference to the following drawings. The elements of the figures are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the application (generally schematic):

[0046] Figure 1 A schematic diagram of an ion spectrometer is shown, with which spectral data can be obtained that have been separated according to a plurality of physico-chemical properties, such as retention time, gas phase ion mobility and mass or mass-to-charge ratio.

[0047] Figure 2 The individual steps in a structure identification workflow applied to spectral data, for example from the apparatus depicted in Figure 1 are schematically shown.

[0048] Figure 3A A match score distribution for one candidate molecular structure is shown.

[0049] Figure 3B A match score distribution for two candidate molecular structures that partially overlap is shown.

[0050] Figure 4 An exemplary plot of spectral match frequency as a function of assumed collision cross section match score is shown, here for a peptide. DETAILED DESCRIPTION

[0051] While the application has been illustrated and described with reference to numerous different embodiments thereof, a skilled person will recognize that various changes can be made in form and detail without departing from the scope of the application as defined in the appended claims.

[0052] Figure 1 A schematic diagram of a possible ion spectrometer apparatus is shown, comprising a plurality of different separation stages with which spectral data can be obtained that are dispersed according to several physico-chemical properties.

[0053] The sample can first be separated in a chromatographic stage, as shown in 2. The chromatographic stage can include a liquid chromatographic stage with a column containing a suitable stationary phase through which the sample dissolved in a suitable mobile phase is flowed. The result can be a series of chromatographic peaks eluting at characteristic retention times later, depending on the settings of the chromatographic conditions.

[0054] The eluent of the chromatographic stage can be passed to an ion source, as shown in 4, which can convert the sample molecules contained in the eluent peaks into gas-borne charged analyte molecules or analyte ions. The ion source can be an electrospray ion source, which utilizes a high voltage difference established at the nozzle with respect to a counter electrode to atomize and ionize a liquid sample, such as a sample eluted from a liquid chromatographic column. Generally, ions can be generated, for example, by using spray ionization (e.g., electrospray (ESI) or thermal spray), desorption ionization (e.g., matrix-assisted laser / desorption ionization (MALDI) or SIMS ionization), chemical ionization (CI), photoionization (PI), electron impact ionization (EI), or gas discharge ionization.

[0055] The analyte ions can be collected and funneled into a well-collimated ion beam, facilitating their efficient transfer to an ion mobility separation stage, as shown in 6. The ion mobility separation stage can utilize the interaction of analyte ions with moving or stagnant gas upon application of an electric field, which is held constant or varies with time. By way of example, the ion mobility separation stage can be designed and configured according to the principles of trapped ion mobility separation (TIMS). Patent publication US 7,838,826 Bl (incorporated by reference in its entirety herein) gives an example of a TIMS separation stage. The result of the ion mobility separation can be a series of eluting ion mobility peaks at characteristic times later, depending on the conditions of the mobility separation device. Other suitable types of gas-phase ion mobility separators can include a drift tube ion mobility separator (DTIMS), a traveling wave ion mobility separator (TWIMS), or a gas-phase ion mobility filter such as a field asymmetric ion mobility separator (FAIMS).

[0056] The eluting mobility peaks can pass through an ion guide stage, as shown in 8, which can be used to pass the analyte ions through a pressure differential between the relatively higher pressure in the ion mobility separation stage and the lower pressure maintained in the subsequent stage for further gas-phase ion handling and operation. Such an ion guide stage can include different ion guides, for example, a multipole rod set ion guide and / or a stacked ring ion guide.

[0057] A filter stage, as indicated at 10, can follow the ion guide stage. The filter stage can comprise a mass filter, for example a quadrupole mass filter, which facilitates transmission of analyte ions in a wideband or precursor screening mode, with the aim of sorting as few ions as possible, or in other words, transmitting as many of the incoming ions as possible, and in one of a bandpass filter mode, a high pass filter mode and a low pass filter mode, with the aim of reducing the transmission window to a relatively narrow mass or mass-to-charge ratio (m / z) range, such that eliminating ions do not fall within the transmission window. The wideband or precursor screening mode and any one of the filter modes can be (fast) continuously alternating.

[0058] A fragmentation stage, as indicated at 12, can follow the filter stage. The fragmentation stage can comprise an ion guide filled with collision gas and further equipped with an electrode which facilitates switching of an acceleration voltage to pull analyte ions at high speed into the collision gas to induce dissociation. Selected precursor ions from the analyte ions can be fragmented into a plurality of characteristic fragment ions. Typically, ions can be fragmented in the fragmentation stage by collision induced dissociation (CID), surface induced dissociation (SID), photo dissociation (PD), electron capture dissociation (ECD), electron transfer dissociation (ETD), collision activation after electron transfer dissociation (ETcD), activation in synchrony with electron transfer dissociation (AI-ETD) or by reaction with highly excited or radical neutral particles.

[0059] Ions emerging from the collision cell can be passed to a mass separation stage, as indicated at 14. The mass separation stage can take the form of a reflectron time-of-flight (rTOF) separation stage, in which orthogonal ions are injected into a time-of-flight flight tube. At the end of a curved flight path within the flight tube, ions can be registered by a collision detector, for example a secondary electron multiplier detector. The result can be a spectrum of ion abundances, for example ion intensities, plotted on a scale related to molecular weight or mass, for example time-of-flight. Together with information from the chromatography stage 2 and the ion mobility separation stage 6, the spectral data can be presented in different plots, for example 3D plots, in which each axis corresponds to a scale of the following physico-chemical properties: (i) retention time from the chromatography stage 2, (ii) gas phase ion mobility or related property from the ion mobility separation stage 6, and (iii) mass or mass-to-charge ratio or related property from the mass separation stage 14, while the abundance of signal peaks can be represented by colour or other suitable graphical features.

[0060] Figure 2 A method of associating molecular structures with signal peaks contained in spectral data obtained from separation according to one or more physico-chemical properties, for example reference Figure 1 is exemplified, is schematically illustrated in several steps.

[0061] First, the obtained spectral data, e.g. one or more signal peaks in the spectrum, is provided, as indicated by 20 on the left side of the figure. The one or more signal peaks can be produced by an ion species of biological origin, e.g. a peptide, a protein, a lipid, a glycan, a polysaccharide, an oligonucleotide, a metabolite, etc. The spectral data can be related to an experimental or experimentally determined value of an ion mobility Km or a related property, e.g. produced by the separation in the ion mobility separation stage 6 in Figure 1 . The related property can for example include a proxy, e.g. a drift time tD of the ion species through the drift tube ion mobility separator, or a derived or deduced parameter, e.g. a collision cross section CCS D,m , (proportional to 1 / K) or a collision cross section to charge ratio (CCS / z) m . m .

[0062] One or more candidate molecular structures can be determined that are suitable to be associated with the one or more signal peaks. The determination can be based on a mass or mass-to-charge ratio filtering during the acquisition of the spectral data, e.g. using a quadrupole mass filter, so as to define a limited mass or mass-to-charge ratio range that the candidates must comply with, e.g. produced by the filter stage 10 in Figure 1 . The one or more signal peaks can have a plurality of experimental values of a second physico-chemical property, and each candidate molecular structure can be associated with one or more candidate values of the second physico-chemical property, as indicated by 22. The one or more experimental values of the second physico-chemical property associated with each candidate molecular structure and the one or more candidate values of the second physico-chemical property can be indicative of a molecular weight of at least one of the precursor ion species and a related fragment ion species of the precursor ion species upon dissociation. Preferably, the second physico-chemical property can include an ion mass m or an ion mass-to-charge ratio m / z. It is also possible that the second physico-chemical property can include a proxy of the ion mass m, e.g. a flight time in a flight tube of a time-of-flight separator.

[0063] In a matching step, the one or more candidate values can show a level of agreement with the one or more experimental values of m / z or the related property, thereby implying a m / z or related property match score SC m / z for each candidate molecular structure, as indicated by 24. Furthermore, the m / z or related property match score SC m / z may be used in a step of associating the molecular structure with the one or more signal peaks. The m / z or related property match score SC m / z may be a scalar and can be computed by a uniform (unity) accumulation each time the m / z or related property value of the one or more signal peaks agrees with or falls into the same m / z or related property set as the candidate value m / z or related property value of the candidate molecular structure being examined. The more signal peaks that match, the higher the m / z or related property match score SCm / z The higher it gets, the more reliable the identification result is. m / z

[0064] The order of the separation according to one or more physico-chemical properties can be set such that the separation according to (gas phase ion) mobility or a related property is at least preceded and followed by the separation according to m / z or a related property. For example, the separation according to m / z or a related property can comprise at least one of mass or mass-to-charge ratio filtering and mass or mass-to-charge ratio dispersion, e.g. time-of-flight dispersion in a flight tube, both exemplarily performed after the separation according to mobility or a related property, as illustrated with the schematic diagram in Figure 1 .

[0065] For each candidate molecular structure, the mobility or related property matching score SC CCS may be provided by one of a calculation, an estimation, a derivation or an extrapolation as a function of the individual distribution of mobility or related property as illustrated in Figure 3A . The calculation, the estimation, the derivation or the extrapolation can comprise a method based on at least one of (i) a statistical evaluation, (ii) a machine learning and (iii) a deep learning of a previously acquired and characterized spectral measurement dataset. Each distribution can be characterized by one or more mobility or related property seed values, which for example define the center or the position of highest score of the distribution. The mobility or related property matching score SC CCS may be indicative of a probability on a scale between a first value (exclusion of match) and a second value (match determination). As illustrated, the first value can be zero and the second value can be one (unity). Each distribution can be configured such that it can lead to mobility or related property matching scores SC CCS deviating from each other. Preferably, each distribution can be configured such that it can lead to a region along the mobility or related property scale with the highest mobility or related property matching score SC CCS , for example indicated at a position close to the center of the distribution, which exhibits a value of 0.97 (first vertical dashed line), and an adjacent region with a reduced mobility or related property matching score SC CCS , for example indicated on the descending flank of the distribution, which exhibits a value of 0.27 (second vertical dashed line). It is clear that there can be more regions of deviating mobility or related property matching scores in the course of the distribution.

[0066] ​Each distribution can follow an analytical function, such as a Gaussian function or other suitable mathematical function, such as a step function or a step-continuous function, which represents a probabilistically weighted deviation or spread of the observed or experimentally determined values of the mobility or related property of the candidate molecular structure under examination. The distribution can also represent an estimation, derivation and / or extrapolation of the mobility or related property match score for a candidate molecular structure for which no prior spectral data exists, in particular by means of deep learning and / or machine learning approaches on existing data sets. In various embodiments, machine learning or deep learning can be performed using a mixture density network (MDN) model.

[0067] The first distribution of the first candidate molecular structure and the second distribution of the second candidate molecular structure can partially overlap, such as in Figure 3B is exemplarily shown. This can mean that the assumed mobility or related property match score SC CCS,p is higher in the case of the first candidate molecular structure than in the case of the other competing candidate molecular structure, on the upper part of the descending flank in the left distribution #1, presenting a value of 0.68, than on the lower part of the ascending flank in the right distribution #2, presenting a value of 0.17, which provides a further consideration factor in the match judgment and allows to improve the match quality. In Figure 3B the example shown, this can be interpreted as the association of one or more signal peaks in the examined spectral data to the candidate molecular structure of distribution #1 is more reliable than the association to the candidate molecular structure of distribution #2.

[0068] As Figure 3A and 3B shown, the assumed mobility or related property match score SC CCS,p of each candidate molecular structure can be defined as the output from the respective distribution upon application or insertion of an experimentally determined value or experimentally determined value of the mobility or related property CCS m (or K m or (CCS / z) m ) of one or more signal peaks. In other words, the assumed mobility or related property match score SC CCS,p may be defined for each candidate molecular structure as a function of the position of the experimentally determined value of the mobility or related property CCS m of one or more signal peaks in the mobility or related property distribution of the candidate molecular structure under examination. This is exemplarily shown in Figure 3A the two vertical dashed lines in the middle, where it is shown that one experimentally determined mobility or related property value falls into the candidate mobility or related property distribution close to the center, resulting in a high confidence assumed match score (0.97), and one experimentally determined mobility or related property value falls into the candidate mobility or related property distribution far from the center, resulting in a lower confidence assumed match score (0.27).

[0069] From Figure 3B Another example can be seen from the single vertical dashed line in FIG. 18, which shows that a single experimental mobility or related property value falls into the first candidate mobility or related property distribution #1 close to the center, resulting in a high confidence, in the left distribution #1 at SC CCS,p = 0.68, while falling into the second competing candidate mobility or related property distribution #2 far from the center, resulting in a low confidence, in the right distribution #2 at SC CCS,p = 0.17. This finding indicates that the candidate molecular structure associated with the left distribution #1 is more likely or more trustworthy match than the candidate molecular structure associated with the right distribution #2.

[0070] Returning to Figure 2 The hypothetical mobility or related property match score SC CCS,p may be used in the step of associating a molecular structure with one or more signal peaks. In a first variant, the hypothetical mobility or related property match score SC CCS,p may be used to exclude a molecular structure from association. In Figure 2 , a low value of the hypothetical mobility or related property match score SC CCS,p may indicate that the potential molecular structure does not match the signal peaks observed in the spectral data 20. A threshold value of the mobility or related property match score SC CCS may be established, which defines on the match score scale a first range to exclude a molecular structure as not applicable, and a second range to accept a molecular structure as possibly real.

[0071] In one embodiment, the hypothetical mobility or related property match score SC CCS,p may be used alone and in combination with the m / z or related property match score SC m / z (and further match or score parameters derived or derived from retention time, intensity of fragment ion species, isotope distribution of ion species, charge state of ion species, etc.) for performing a (quadratic) discriminant analysis to distinguish between trustworthy matches and untrustworthy matches.

[0072] In a further embodiment, the hypothetical mobility or related property match score SC CCS,p and the m / z or related property match score SC m / z may be combined to generate a third match score as shown in 28, and the third match score can be used in the step of associating a molecular structure with one or more signal peaks. The combination can include the hypothetical mobility or related property match score SC CCS,p and the m / z or related property match score SC m / zmultiplication or other suitable mathematical operation, as indicated at 30. The candidate molecular structure having the most extreme value of at least one of the mobility or related property match score SC CCS , the m / z or related property match score SC m / z The candidate molecular structure having the most extreme value of at least one of the first, second, and third match scores can be used to associate the molecular structure with one or more signal peaks. Preferably, the candidate molecular structure having the highest match score of at least one of the first, second, and third match scores can be used to associate the molecular structure with one or more signal peaks. m / z , the assumed mobility or related property match score SC CCS,p The candidate molecular structure having the highest match score of at least one of the first, second, and third match scores can indicate that the association of one or more signal peaks with the candidate molecular structure is likely true.

[0073] One or more candidate molecular structures can be determined from a pool of target candidates indicative of possible molecular structures and a pool of decoy candidates indicative of impossible molecular structures, as indicated at 32 in Figure 2 The assumed mobility or related property match score SC CCS,p may be used to define a metric that helps to distinguish between trustworthy associations and untrustworthy associations. This additional information, derived inter alia from ion mobility separation, false discovery rate (FDR) or decoy hit rate (DHR) calculations, can be used to increase accuracy.

[0074] Figure 4 An exemplary illustration is shown in a graph showing the peptide spectrum match (PSM) frequency on the vertical axis (y-axis) as a function of the assumed collisional cross section match score SC CCS,p , involving both target candidate matches (solid columns) and decoy candidate matches (hollow columns). It can be seen that at high match scores close to 1, the target candidate matches dominate, while at low match scores the frequency is almost evenly distributed between target candidate matches and decoy candidate matches. It is well known that decoy candidate matches cannot be true, while there is always at least some uncertainty whether a target candidate match is true or not, especially in automated processing where no experienced practitioner carefully checks the results. The information from decoy candidate matches can be used to define in a statistical approach a landmark on the match score scale that defines a first range of match scores that are considered untrustworthy due to low confidence, whether the candidate match is a decoy or a target candidate match, and a second range where the candidate match can be considered likely or most likely to be true.

[0075] This is exemplified by Figure 4The vertical solid line in the middle indicates a match score of about 0.5. To the left of this landmark (≤ 0.5), the match can be considered untrustworthy, while to the right of this landmark (> 0.5), the match can be assumed to be trustworthy or likely to be true. The landmark can be chosen such that only a certain low percentage of decoy candidate matches lie within the range indicating trustworthiness (the correct range, higher scores), e.g. five percent of the total number of matches, or less. This additional migration-related measure can advantageously be used to calculate the false discovery rate or the decoy hit rate, for example.

[0076] The application has been illustrated and described above with reference to a number of different embodiments of the application. However, those skilled in the art will understand that various changes in the details of the embodiments of the application, or in the application itself, can be made without departing from the scope of the application, if any, and that individual aspects or details of different embodiments can be interchanged or combined, if and as appropriate, in order to form additional embodiments not presently described. In general, the foregoing description has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It is intended that the scope of the application be limited not with this detailed description, but rather by the claims appended hereto, appropriately interpreted in accordance with the doctrine of equivalents, if and as appropriate.

Claims

1. A method for associating molecular structures with signal peaks in spectral data obtained from separation according to one or more physicochemical properties, the method comprising: - providing one or more signal peaks in the acquired spectral data that correlate with experimental values ​​of mobility or mobility-related properties; - determining one or more candidate molecular structures suitable for associating with the one or more signal peaks; - providing a first distribution of matching scores as a function of mobility or a mobility-related property for each candidate molecular structure using one of calculation, estimation, deduction, and inference; - defining for each candidate molecular structure a putative first match score as an output from the corresponding distribution when applying experimental values ​​of the mobility or mobility-related property of the one or more signal peaks; and - Using a putative first match score in the step of associating a molecular structure with one or more signal peaks. The method of claim 1 , wherein the putative first matching score is used to exclude molecular structures from association.

3. A method according to claim 1 or 2, wherein one or more signal peaks have one or more experimental values ​​of a second physicochemical property, and each candidate molecular structure is associated with one or more candidate values ​​of the second physicochemical property, and the one or more candidate values ​​show a level of consistency with the one or more experimental values ​​of the second physicochemical property, thereby implying a second matching score for each candidate molecular structure, and the method also includes using the second matching score in the step of associating the molecular structure with the one or more signal peaks.

4. The method of claim 3, further comprising combining the putative first match score and the second match score to generate a third match score, and using the third match score in the step of associating the molecular structure with the one or more signal peaks. 5 . The method of claim 4 , wherein the candidate molecular structure having the most extreme value of at least one of the putative first match score, the second match score, and the third match score is used to associate the molecular structure with the one or more signal peaks.

6. A method according to claim 3, wherein one or more experimental values ​​of the second physicochemical property associated with each candidate molecular structure and one or more candidate values ​​of the second physicochemical property indicate the molecular weight of at least one of the precursor ion species and the associated fragment ion species of the precursor ion species upon dissociation.

7. The method according to claim 3, wherein the separation based on mobility or a mobility-related property is at least preceded or followed by the separation based on the second physicochemical property.

8. The method of claim 3, wherein separation according to the second physicochemical property comprises at least one of mass or mass-to-charge ratio filtering and mass or mass-to-charge ratio dispersion.

9. The method of claim 8, wherein separation according to the second physicochemical property comprises time-of-flight dispersion in a flight tube.

10. The method of claim 1, wherein each distribution is configurable such that it results in first matching scores that deviate from one another.

11. The method of claim 1 , wherein each distribution is configurable such that it results, along a mobility or mobility-related property scale, in a region having a highest first match score and adjacent regions thereto having decreasing first match scores. The method of claim 11 , wherein the first distribution of the first candidate molecular structure and the second distribution of the second candidate molecular structure partially overlap.

13. The method of claim 1, wherein the first match score indicates a probability on a scale between a first value and a second value, wherein the first value indicates a match is ruled out and the second value indicates a match is confirmed.

14. The method of claim 1, wherein calculating, estimating, inferring, or extrapolating comprises a method based on at least one of (i) statistical evaluation, (ii) machine learning, and (iii) deep learning of a previously acquired and characterized spectroscopic measurement data set.

15. The method of claim 1, wherein the one or more signal peaks are generated by ionic species of biomolecular origin.

16. The method of claim 1 , further comprising determining one or more candidate molecular structures from a target candidate pool indicating possible molecular structures and a decoy candidate pool indicating impossible molecular structures, and using the assumed first match score to define a metric that helps distinguish between credible associations and implausible associations.

17. A method for associating molecular structure with signal peaks in spectral data obtained from a separation based on one or more physicochemical properties, the method comprising: - providing a plurality of signal peak groups and a plurality of experimental values ​​of mobilities or mobility-related properties in the acquired spectral data, each signal peak group being associated with an experimental value of the mobility or mobility-related property and having one or more signal peaks; - determining a plurality of candidate molecular structure groups from a target candidate pool indicating possible molecular structures and a decoy candidate pool indicating impossible molecular structures, each candidate molecular structure group having one or more candidate molecular structures and being suitable for being associated with one or more signal peak groups; - providing, for each candidate molecular structure, one or more candidate values ​​of mobility or mobility-related properties by one of calculation, estimation, deduction, and inference; - providing a plurality of match scores by defining the one or more match scores for each candidate molecular structure as a function of the level of agreement between the one or more candidate values ​​of the mobility or mobility-related property and the plurality of experimental values ​​of the mobility or mobility-related property for the plurality of signal peak groups, and - Use multiple matching scores to define metrics that help distinguish between credible and uncredible associations. The method of claim 17 , wherein the matching score is a scalar.

19. The method according to claim 17 or 18, further comprising establishing match score landmarks on a match score scale that define a first range that is assumed to indicate an untrustworthy association, regardless of whether the potential candidate molecular structure is from a decoy candidate pool or a target candidate pool, and a second range that is assumed to indicate a trustworthy association.

20. The method of claim 19, wherein the match score landmark is defined as a percentage within a second range that is less than one of five percent, four percent, three percent, two percent, and one percent of the set of signal peaks found to be associated with candidate molecular structures in the decoy candidate pool.

21. Apparatus for recording ion species separated according to one or more physicochemical properties, said apparatus comprising a data processing unit designed and configured to perform the method according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Apparatus and method for parallel flow ion mobility spectrometry combined with mass spectrometry

    US7838826B1

  • System and method for scoring peptide matches

    WO2004013635A2

  • A method and system of estimating the cross-sectional area of a molecule for use in the prediction of ion mobility

    WO2010119289A2

  • A method and system of estimating the cross-sectional area of a molecule for use in the prediction of ion mobility

    WO2010119293A2

  • A method of screening a sample for the presence of one or more known compounds of interest and a mass spectrometer performing this method

    WO2011027131A1