Methods for collision cross-section prediction and ion mobility tandem ms analytical methods using such methods

By training machine learning models on structurally similar molecular structures using fingerprint similarity, the method addresses the reliability issues of existing CCS prediction tools, enhancing prediction accuracy and annotation confidence in complex samples.

WO2025247796A1PCT designated stage Publication Date: 2025-12-04BRUKER DALTONIK GMBH & CO KG
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/064433
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-29
Filing Date
2025-05-26
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing collision cross-section (CCS) prediction tools for ion mobility measurements in complex samples are insufficiently reliable, particularly in bioanalytical workflows like proteomics, lipidomics, and metabolomics, due to inaccurate generalization or reliance on prior knowledge of compound classes, leading to reduced specificity and prediction accuracy.

Method used

A method involving machine learning models trained on structurally similar molecular structures, using fingerprint similarity to select training data, eliminating the need for semantic classification, and ensuring homogeneous training sets for improved prediction accuracy.

Benefits of technology

Enhances the precision and generalizability of CCS predictions by aligning training data with molecular features, providing objective and reproducible results without manual classification, thus improving annotation confidence in complex samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025064433_04122025_PF_FP_ABST
    Figure EP2025064433_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A computer implemented method of training machine learning models for the predictive calculation of a collision cross-section (CCS) value of a molecular structure belonging to a similarity structure family, comprising: collecting data from an existing database with at least molecular structure information and associated values of collision cross-section values, in that for training for a model for that similarity structure family a training subset from that database is generated for said similarity structure family, in that a) molecular fingerprints are calculated for the molecular structures in the database, and b) the calculated molecular fingerprints are clustered in similarity groups of (highest) fingerprint similarity, c) and for training the machine learning models for prediction of a similarity structure family only the information of molecular structures belonging to the same similarity group is used.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TITLE

[0002] METHODS FOR COLLISION CROSS-SECTION PREDICTION AND ION MOBILITY TANDEM MS ANALYTICAL METHODS USING SUCH METHODS

[0003] TECHNICAL FIELD

[0004] The present invention relates to methods for predicting collision cross-section (CCS) values for ion mobility measurements, in particular involving machine learning algorithms, and corresponding analytical methods, in particular LC coupled ion mobility tandem MS analytical methods, making use of such predictions.

[0005] PRIOR ART

[0006] US 7,838,826 B1 (M. A. Park, 2008) and the corresponding patent family members presents a small ion mobility analyzer / spectrometer which has become known under the acronym “TIMS” analyzer / spectrometer (TIMS = trapped ion mobility spectrometry). The terms ion mobility analyzer and ion mobility spectrometer are used interchangeably here. A TIMS analyzer comprises a gas flow that drives ions against a counter-acting electric field barrier such that the ions are at first trapped along the axis of the TIMS analyzer. The ions are confined in the radial direction by an electric RF field. After transferring ions from an ion source to the electric field barrier, the height of the electric field barrier or the gas velocity is adjusted such that ion species are released from the electric field barrier in the sequence of their mobility.

[0007] Commonly, the length of the ion mobility separation unit of a TIMS analyzer amounts to about five centimeters only. In a small tube with an inner diameter of about eight millimeters, a radial RF quadrupole field is generated to hold ions near to the axis. A gas flow inside a tube drives ions entrained in the gas flow against a ramped counter-acting electric DC field barrier where the ions are trapped and separated according to their mobilities, which depend on the CCS of the precursor ions, at locations on the field ramp at which the friction force of the moving gas equals the counter-acting force of the electric DC field on the ramp. After loading the TIMS with ions, the height of the electric DC field barrier is decreased; this scan releases the ion species in the sequence of their mobility. Unlike many other trials to build small ion mobility spectrometers, the small device by M. A. Park has already achieved, with reduced scan speeds, ion mobility resolutions up to RmOb = 400, which is extraordinarily high.

[0008] A TIMS analyzer with parallel ion accumulation is described in US 9,683,964; it improves the utilization of the ions from the ion source to nearly 100%. Importantly, many ion sources, such as electrospray ion sources produce ions continuously. TIMS with parallel accumulation can also be operated to collect and analyze ions continuously - that is, TIMS can operate at 100% duty cycle. Barring pseudopotential or space charge effects, substantially all ions of the ion source are collected and analyzed without loss. TIMS with parallel ion accumulation further provides the possibility to prolong the ion accumulation and scan duration, thereby increasing the ion mobility resolution so as to separate and detect more ion species.

[0009] The ions are collected in an accumulator unit, preferably almost identical to the scanning unit, at a ramp of an electric DC field barrier such that they get spatially separated by their ion mobility along the ramp. Therefore, the accumulated ions are less influenced by space charge than in other types of accumulator units. Of greatest importance, however, is the unique feature of a TIMS analyzer that a longer accumulation period permits to increase the mobility resolution by choosing correspondingly longer mobility scan durations, e.g. 100 milliseconds scan duration with an ion mobility resolution of RmOb = 75 instead of 20 milliseconds scan duration with RmOb = 30. As a consequence of the higher number of ions collected and the better ion mobility resolution, more ion species can be detected and measured. Once an ion mobility scan is completed (optionally after twenty to some hundred milliseconds), the accumulated ions are transferred (in about a millisecond) from the accumulation unit to the scanning unit, and the next ion mobility scan can be started. In total, a skilled practitioner will appreciate that it will be possible to achieve a measurement rate of 300 to 450 ion species per second. If TIMS with parallel ion accumulation is installed in tandem mass spectrometer (MS / MS instrument) an MS-MS instrument, 300 to 450 characteristic fragment ion spectra per second may be measured quantitatively.

[0010] Some improvements for higher amounts of stored ions in selected regions of ion mobility, particularly for ions of low ion mobility, are given in US 9,304,106 B1. The higher loading capacity is based on non-linear electric DC field ramps, with flatter field ramps for ion species of interest, in order to diminish the effect of space charge for these ion species. But for precise ion mobility analyses of low abundant ion species in complex mixtures the influence of the space charge is still high.

[0011] TIMS extends conventional liquid chromatography-mass spectrometry (LC-MS) and bioanalytical workflows, such as proteomic, lipidomic, metabolomic, drug metabolism, structure elucidation or other methods, with an additional ion mobility dimension. In addition to the general benefits of the additional signal separation, the implementation of TIMS in the Bruker timsTOF family of instruments has demonstrated that CCS values of analytes, i.e. peptides, lipids, or metabolites, are highly reproducible. The additional separation dimension already allows in data dependent acquisition (DDA) methods to identify precursors of isomeric compounds to be scheduled for fragmentation individually. As omics workflows aim at covering as much of the chemical complexity of a sample with tandem- MS spectra by DDA, TIMS can already greatly improve the analytical coverage of complex samples.

[0012] WO-A-2019096852 relates to use of an isobaric label in mass spectrometry (MS) analysis using data-independent acquisition (DIA), wherein said isobaric label comprises or consists of a group which fragments in the mass spectrometer (i) at an energy below the energy required for fragmenting analyte-derived precursor ions and / or a higher conversion rate than said precursor ions; and (ii) at said energy according to (i) and when coupled to a precursor ion, at a single site within said group, to yield a first moiety and a second moiety, said second moiety being coupled to said precursor ion. It proposes the use of a trapped ion mobility spectrometry- time of flight (timsTOF) instrument, equipped with parallel / serial fragmentation (PASEF); see, e.g. Meier et al. 2015, doi: 10.1021 / acs.jproteome.5b00932. US-A-2019371585 relates to selection of precursors from a measured mobility-mass map for tandem mass spectrometry and is based on processing a peak list from measured signals and clustering these peaks in the mobility-mass space.

[0013] US-A-2022034840 discloses an apparatus and a method of data independent combined ion mobility and mass spectroscopy analysis which includes introducing precursor ions into an ion mobility spectrometer (IMS), sequentially releasing precursor ions from said IMS according to their ion mobility, introducing said released precursor ions into a mass filter, fragmenting the precursor ions transmitted through said mass filter to generate fragment ions, and carrying out a mass spectroscopy measurement on said fragment ions. The IMS and mass filter are controlled in a synchronized manner to carry out a plurality of IM scans, wherein adjacent mass windows in said IM scan that are associated with consecutive mass spectroscopy measurements of fragment ions overlap, such that precursor ions transmitted through said mass filter during said IM scan are located in at least one continuous scan region in an m / z-IM plane which extends in a generally diagonal direction in said m / z-IM plane.

[0014] Targeted selection of peptides from complex protein digests subjected to fragmentation is a performance-limiting factor in tandem mass spectrometry (MS). The problem is mainly attributed to the large number of co-eluting high abundance peptides and limitations in the duty cycle of the mass spectrometer arranged to operate in data-dependent acquisition (DDA) mode whereby precursor ions are selected sequentially for fragmentation in order of decreasing intensity. Furthermore, the depth of proteomic analyses is also limited by the overwhelming proportion of high abundance background ions that compromise the identification and quantification of low abundance peptides. Whether low-intensity precursors are selected for fragmentation is primarily dictated by the speed at which the mass analyzer can perform tandem MS analysis. The duty cycle of the MS remains a problem even in data-independent acquisition (DIA) mode where all ions within a selected mass-to-charge range are selected for fragmentation. The issue of sample complexity is routinely addressed by fractionating the sample.

[0015] Liquid chromatography coupled to Mass Spectrometry (LC-MS) has now been used for many years in the proteomic community for the identification and quantification of peptides (and thus proteins) from complex sample mixtures. In proteomics, the analytes are typically peptides generated by tryptic digestion of protein samples. The commonly most used approaches are variants of the so-called LC-MS / MS or “shotgun” MS approach that is based on the generation of fragment ions from precursor ions that are automatically selected based on the precursor ion profiles (data dependent analysis, DDA).

[0016] The most mature technology is called Selected Reaction Monitoring (SRM), frequently also referred to as Multiple Reaction Monitoring (MRM). The targets for MRM experiments are defined on a rational basis and depend on the hypothesis to be tested in the experiment. Selected combinations of precursor ions and fragment ions (so called transitions, the set of transitions for one target precursor is called MRM assays) for these targets are programmed into a mass spectrometer, which then generates measurement data only for the defined targets. Thus, in SRM, all transitions are monitored one at a time.

[0017] Yang et al. in Molecules 2022, 27(19), 6424; htps: / / doi.org / 10.3390 / molecules27196424, report that high-resolution mass spectrometry is a promising technique in non-target screening (NTS) to monitor contaminants of emerging concern in complex samples. Current chemical identification strategies in NTS experiments typically depend on spectral libraries, chemical databases, and in silico fragmentation tools. Small molecule identification remains challenging due to the lack of orthogonal sources of information (e.g., unique fragments). Collision cross section (CCS) values measured by ion mobility spectrometry (IMS) offer an additional identification dimension to increase the confidence level. Thanks to the advances in analytical instrumentation, an increasing application of IMS hybrid with high-resolution mass spectrometry (HRMS) in NTS has been reported in the recent decades. Several CCS prediction tools have been developed. However, limited CCS prediction methods were based on a large scale of chemical classes and cross-platform CCS measurements. They successfully developed two prediction models using a random forest machine learning algorithm. One of the approaches was based on chemicals’ super classes; the other model was direct CCS prediction using molecular fingerprint. Over 13,324 CCS values from six different laboratories and PubChem using a variety of ion-mobility separation techniques were used for training and testing the models. The test accuracy for all the prediction models was over 0.85, and the median of relative residual was around 2.2%. The models can be applied to different IMS platforms to eliminate false positives in small molecule identification.

[0018] SUMMARY OF THE INVENTION

[0019] The CCS values are important for confident compound annotation in an analytical workflow involving liquid chromatography-mass spectrometry (LC-MS) with ion mobility dimension. However, very often, if not in most cases, CCS information is not available for each and every precursor that could be present in a complex mixture.

[0020] This severely hampers the possibility of using automatic annotation workflows in such analytical workflows to enable annotation with large databases and libraries.

[0021] This can be addressed by using CCS prediction to close this gap of lacking reference CCS values. Also, machine learning can be used in this context, wherein the accuracy of the corresponding predictions made by the resulting machine learning routines heavily depends on the training data.

[0022] The CCS prediction tools can be trained on either broad sets of small molecules or they can be trained for distinct compound classes, for example lipids, steroids, and the like.

[0023] In the former case, so where the training is based on small molecule data, the CCS prediction tools need to generalize, which hampers their accuracy significantly, with many false or rather offset or inaccurate predictions. There is correspondingly a significant loss in specificity and prediction accuracy in case of small molecule training database models.

[0024] In the latter case, the CCS prediction models are only applicable for the respective compound class, which in turn requires prior knowledge of the compound classes of the input substances, which is normally not the case. The problem with compound class specific CCS prediction tools is further that members of the same compound class normally indeed do share distinct molecular features, but are not necessarily structurally similar. So, fragments or digests may be formed comprising CCS determining structural elements, which are not common elements of the corresponding compound class, and for these the models are highly unreliable.

[0025] As a result, this means that the existing CCS prediction tools are insufficiently reliable and are not really closing the gap of lacking reference CCS values in practice for complex samples or mixtures, so in particular for bioanalytical workflows, such as proteomic, lipidomic, metabolomic, drug metabolism, structure elucidation or other methods.

[0026] The present invention provides for an improved and objective CCS prediction approach, which involves at least one of the following elements: individual models for structural similarity families are trained and used then for prediction; the structural similarity to be within one of these models is established by using fingerprinting of the molecular structures and calculating similarity values between individual molecular structures, and for training for a model for one structural similarity family, only information from databases of molecular structures with known CCS values having a minimum similarity is used; so the proposed approach in particular exclusively relies on, for training of a model for a given similarity structure family, selecting data from a pool of available data based on calculated structural similarity criteria and without taking account of generic (semantic) super-classes such as lipids and lipid-like molecules, benzenoids, organic acids and derivatives, organic oxygen compounds, organoheterocyclic compounds, or the like, as a first of additional data selection criterion for the training of that model.

[0027] Generally speaking, the proposed method, for the selection of training data for training individual models for structural similarity families (or for selecting an individual model for a CCS prediction of a given structure) works with a pure and automated structural similarity approach, by fingerprinting of the molecular structures and calculating similarity values between individual molecular structures, and does not make use of any semantic or manually defined classification for the data selection for training or for the model selection, for the final prediction of the CCS value of a specific structure with unknown CCS value, for that specific structure its fingerprint is calculated, and then similarity values between at least elements of the training data or all the elements of the training data or representative elements of the training data for the trained several structural similarity family models are calculated based on fingerprints, and the model (or the models) of the structural similarity models with the highest similarity with the fingerprint of the specific structure is (are) used for the predictive calculation of the CCS value of that specific structure. a) In a first step, separate models for different “similarity families” are created. a. Instead of “compound classes” in the sense of generic super-classes, “similarity families” are created from structures in the overall set of training data without semantic meaning of compound classes or ontologies, but based on objectively calculated (similarity) properties of structures, such as by creating structure fingerprints. b. These “similarity families” do not have to be transparent to the users. c. For each similarity family, a dedicated CCS prediction model is trained and optimized. For this model, only data sets from the overall data pool are used, which fall within an objectively calculated similarity score for or around that similarity family. b) To predict the CCS value of new structures (target structures), the most suitable model(s) are first determined and then applied. a. At first, for the target structure the same objective (structural pattern related) properties are calculated. b. According to these properties, the most suitable model or multiple suitable models are selected. The most suitable model(s) are the one(s) where the training structures or data sets are most similar to the target structure using an objectively calculated similarity score. The best model(s) can be determined by a similarity of the fingerprint of the target compound to the most similar fingerprints of training compounds in the respective models. One preferred approach is using the Representative Structure Similarity (RSS), that preferably represents the average of the top two, three, four or five Tanimoto coefficients the query structure scores with structures in the training data of the model. See also Zhou et al in "Ion mobility collision crosssection atlas for known and unknown metabolite annotation in untargeted metabolomics", Nat Commun i s, 4334 (2020). https: / / doi.org / 10.1038 / s41467-020-18171-8, including the supplementary material thereof, which is included into this disclosure a concerns the RSS values and the uses thereof for similarity assessment. c) The most suitable model(s), so the ones having the largest similarity of the underlying training data with the target structure, are used to predict a CCS value for the target structure. a. If multiple similarity models are applied, the outputs may be consolidated to one final predicted CCS value. i. by creating the mean ii. by creating the median iii. by creating a weighted mean (based on RSS between target structure and model) d) The presented workflow cannot only be provided and applied as a standalone tool, but can also be fully integrated into a data processing pipeline from raw data processing, over statistical analysis, and compound annotation. Integrated into the compound annotation of such a workflow, it provides automated and on-the-fly improvement to annotation confidence in a large scale. e) The output of the prediction can finally be matched against the measured CCS value of candidate features in e.g. complex samples, to confirm or reject putative annotations. The absolute or relative deviation of predicted vs. measured CCS values may also be used to create an annotation confidence score. Ideally, this annotation confidence score also incorporates other criteria, like mass deviation, isotopic pattern similarity and MSMS spectral similarity.

[0028] According to a first aspect of the present invention, it relates to a computer implemented method of training machine learning models for the predictive calculation of a collision cross-section (CCS) value of a molecular structure belonging to a similarity structure family. The proposed method comprises the following steps: collecting data from an existing database with at least molecular structure information and associated values of (measured) collision cross-section values, in that for training for a model for that similarity structure family a training data subset from that database is generated for said similarity structure family.

[0029] That subset is generated in that a) molecular fingerprints (of the molecular structures) are calculated for the molecular structures in the database, and b) the calculated molecular fingerprints are clustered in similarity groups of fingerprint similarity, preferably highest fingerprint similarity, c) and for training the machine learning model for prediction of a similarity structure family, only the information of molecular structures belonging to the same similarity group is used.

[0030] The data selection process for the training of a model for such a cluster as determined in step b), which then represents a similarity structure family, is exclusively based on the molecular fingerprints and is not additionally or previously supplemented by other selection criteria such as whether data sets belong to super-classes or not.

[0031] While previous approaches — such as those described in Yang et al. (Molecules 2022, 27(19), 6424) have made significant contributions to the field by introducing semantic superclass-based data selection strategies, our method offers a distinct and complementary advancement by relying exclusively on structural similarity criteria derived from molecular fingerprint calculations.

[0032] In our approach, the selection of training data for a specific predictive model is entirely automated and based solely on quantifiable structural similarity. This eliminates the need for manual classification into semantic superclasses or reliance on domain-specific prior knowledge. As a result, the formation of similarity structure families — and the assignment of predictive models to these families — is both objective and reproducible. Importantly, this process is fully transparent to the user: the only required input is the target molecular structure, after which the system autonomously identifies the most appropriate similarity family and applies the corresponding model to predict the CCS value.

[0033] Through comparative analysis, we observed that superclass-based selection methods, while conceptually intuitive, can inadvertently introduce structural heterogeneity into training datasets. This can lead to reduced model accuracy due to the inclusion of structurally dissimilar compounds or the exclusion of relevant ones that fall outside predefined semantic categories. Whether superclass filtering is applied before, after, or in parallel with structural similarity assessments, it tends to either dilute the structural coherence of the training set or unnecessarily constrain it — both of which can impair model performance.

[0034] By contrast, our structurally driven approach ensures that training data is consistently aligned with the molecular features most relevant to the prediction task, thereby enhancing model precision and generalizability.

[0035] According to a preferred embodiment of this method, the similarity groups of (highest) fingerprint similarity are formed / clustered around chemical substructure units. This is what automatically happens in the above-mentioned step b).

[0036] It is noted that when generating the clusters of similarity groups, these clusters can be formed in an exclusive way, such that information is only used in one cluster. However, it may also be advantageous to form overlapping clusters, and in particular in that case it may be advantageous, for the prediction, to use models which are derived from overlapping clusters.

[0037] The present invention furthermore relates to a computer implemented method for the predictive calculation of a collision cross-section value of a specific molecular structure (of which the collision cross-section is unknown), wherein for the calculation of the collision cross-section of that specific molecular structure a) in a first step the molecular fingerprint of that specific molecular structure is calculated; b) in a second step the fingerprint similarity between the fingerprint of that specific molecular structure calculated in a) and at least one, preferably at least 10, or at least 100, or all structures of the training data for several similarity structure family prediction models is calculated; c) for the predictive calculation of the collision cross-section of that specific molecular structure the model for the similarity structure family having the highest fingerprint similarity established in step b) is used, or models of several similarity structure families with the highest fingerprint similarity established in step b) are used for calculation and used for prediction and the final predicted collision cross-section value is taken as an average, weighted average or median value of these calculated predictive values.

[0038] According to a first preferred embodiment of such a prediction method, in step c) a model obtained in a similarity model training method as described above is used.

[0039] The molecular fingerprint is preferably a binary substructure fingerprint in the form of an ordered list of binary bits, wherein each bit represents a Boolean determination of or test for the presence of elements in the structure, preferably of at least one of an element count, an element type, connection type, environment, in the chemical structure.

[0040] Preferably the substructure fingerprint is according to the PubChem standard (V1.3 http: / / pubchem.ncbi.nlm.nih.gov).

[0041] The fingerprint similarity in step b) is preferably established using the Tanimoto coefficient. The Tanimoto coefficient Simian (S1 ,S2) between two structures S1 and S2 is determined as follows:

[0042] Simian (S1 ,S2) is in the range of [0.0, 1.0], S1S2 stands for the evaluation of all elements present in both S1 and S2, S10niyand S2oniystand for the evaluation of elements present in one of the structures S1 and S2, respectively, only.

[0043] An example is as follows:

[0044] The fingerprint similarity in step b) can be established using the highest Tanimoto coefficient between the specific molecular structure and the structures used to train the respective machine learning model.

[0045] The fingerprint similarity in step b) may also be established using the average of the two, three, five or ten best Tanimoto coefficients between the specific molecular structure and the structures used to train the respective machine learning model.

[0046] The fingerprint similarity in step b) of the training method is preferably established using the Tanimoto coefficient, and wherein for clustering, structures having a similarity of at least 0.8, or at least 0.9 or at least 0.95 are used.

[0047] Furthermore, the present invention relates to an analytical method using the above- mentioned technologies.

[0048] More specifically, it relates to a method of data dependent or data independent combined liquid chromatography (LC), ion mobility and tandem mass spectroscopy analysis, comprising the following steps: introducing precursor ions resulting from LC separation into at least one trapped ion mobility spectrometry (TIMS) separator, and separating the precursor ions according to mobility in the trapped ion mobility spectrometry (TIMS) separator, sequentially releasing precursor ions from said trapped ion mobility spectrometry (TIMS) separator according to their ion mobility, introducing said released precursor ions into a mass filter which selectively transmits precursor ions having m / z values falling within a controllable mass window, fragmenting the precursor ions transmitted through said mass filter to generate fragment ions, carrying out a mass spectroscopy measurement on said fragment ions, wherein each fragment ion is associated with a mass window and an ion mobility (IM) range, and associating detected fragments with its corresponding precursor ion, wherein said trapped ion mobility spectrometry (TIMS) separator and said mass filter are scheduled in a synchronized manner such as to carry out a plurality of ion mobility (IM) scans, during which precursor ions of increasing or decreasing IM are successively released from said second ion mobility separator (IMS), and during which the mass window of said mass filter is shifted continuously or stepwisely towards lower or higher m / z values, respectively, wherein for annotation of signals, collision cross-section values of a specific molecular structure from databases not including collision cross-section information, cross-section value prediction for that specific molecular structure as described above is used.

[0049] At the beginning of one LC observation retention time window preferably a survey scan is taken, wherein in that survey scan a full ion mobility width of interest and a full m / z width of interest is scanned, and wherein from that survey scan automatically a list of detectable signals is determined, from that list of detectable signals the strongest detectable signals are automatically selected for scheduling at least once in a synchronized manner in that LC observation retention time window.

[0050] Last but not least, the present invention relates to the use of the methods for the analysis of a complex sample, including samples in proteomics, lipidomics, and metabolomics. A complex sample is a mixture of many different components in variable or inexact proportions, usually natural or natural based, such as plant or animal extracts, or derivatives (e g. digests such as protease digests) thereof.

[0051] Further embodiments of the invention are laid down in the dependent claims.

[0052] BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Preferred embodiments of the invention are described in the following with reference to the drawings, which are for the purpose of illustrating the present preferred embodiments of the invention and not for the purpose of limiting the same. In the drawings,

[0054] Fig. 1 shows an overview over the structure similarity family training approach;

[0055] Fig. 2 shows an overview over the prediction using such a structure similarity family model; Fig. 3 shows an example of a classification based on superclasses leading to less structurally homogenous training data selection.

[0056] DESCRIPTION OF PREFERRED EMBODIMENTS

[0057] To give an example of the current problem and the solution, the following: Assume a user has acquired LC-TIMS-PASEF data of complex samples to investigate metabolic differences in healthy and diseased patients. Some of the detected signals show significant differences in their intensities, when comparing healthy and diseased patients. To link this finding to biological understanding, it now is important to identify (“annotate”) the metabolites, which create these signals. For a set of known compounds, which can be acquired as reference standards, this annotation is performed by matching the signals m / z, RT, isotope pattern, MSMS spectrum and CCS value in the complex sample to measurements of the respective standards. But not all compounds can be acquired as references, but only exist in form structural information in libraries. The respective m / z and isotope patterns can be calculated from structures. To match measured MSMS spectra against structures several solutions for in-silico fragmentation exist.

[0058] For the prediction of CCS values from structural information, several CCS prediction tools are possible.

[0059] However, these prediction tools are either class-specific; that means one can only apply them to a certain class of compounds AND one must know the respective class of the compound (which may be ambiguous) - or they are general models, which are said to be applicable to all classes of small molecules, but then strongly generalize and are thus less specific for certain compound classes (or individual compounds).

[0060] This demands for a strategy, that provides the accuracy / specificity of class specific models, without putting the burden of compound classification and finding and selection of the most suitable model onto the user.

[0061] Not for all structures, CCS reference values exist. Still, CCS matching is an important contributor to annotation confidence. Machine-learning based predictions of CCS values for new structures are possible, but the accuracy of the predictions depends on the similarity of the new structure to the structures present in the training data of the used prediction models. While prediction models typically get better with increasing amounts of training data, the problem of generalization remains. And even if the model generalizes well, it in turn may lack specificity for certain classes of molecules.

[0062] Users however want to apply and trust CCS prediction models without prior technical knowledge. Users do not want to have to preselect the correct prediction model for each compound at hand. While this may be feasible for a broad class such as lipids, for other classes of compounds the step of classification itself is difficult and ambiguous.

[0063] Machine learning based CCS prediction models can be specific for certain compound classes (such as lipids), or, when they cover a broader range of small molecules, are less accurate.

[0064] While e.g. lipids are a large class of compounds, that due to their common way of composition, if no derivatives are included, can be quite comprehensibly described with head group, chain length and double bonds, which in turn largely correlate with CCS values, other classes of compounds are much more heterogeneous and thus are harder to describe with only a few descriptors. Consequently, it is not feasible to generate accurate CCS prediction models for all classes of compounds. To select the correct compound-class- specific model to predict the CCS value of a new structure, the respective compound-class needs to be determined.

[0065] Fig. 1 schematically illustrates how the molecular structures for training are fingerprinted, subsequently clustered for similar structures based on fingerprint similarity, and how the corresponding models for prediction are generated.

[0066] Fig. 2 schematically illustrates how a specific input structure, of which the CCS value is not known, is first fingerprinted, then the fingerprint similarity (e.g. RSS) is established to the corresponding training data of the respective similarity models, and subsequently the closest model, or the closest models, for the similarity families are used for prediction.

[0067] As previously discussed, the accuracy of predictive modeling is significantly enhanced when both the formation of structural similarity groups and the selection of training data are based solely on calculated similarity metrics — specifically, molecular fingerprint-based similarity scores such as the Tanimoto coefficient. Importantly, our methodology avoids the use of semantic or manually defined classification criteria, ensuring that the grouping and model selection processes remain objective and data-driven.

[0068] Equally critical is the model selection step for a given target structure. In our approach, this is determined exclusively by the structural similarity between the target and the training data (or subsets thereof) associated with each model. This ensures that the most structurally relevant model is applied, thereby improving prediction reliability.

[0069] To illustrate the limitations of semantic superclass-based grouping, we conducted a comparative analysis using the classification data from Yang et al. (Molecules 2022, 27(19), 6424). While the work by Yang et al. represents a valuable contribution to the field, our analysis revealed that the training sets within individual superclasses are not always structurally homogeneous. Specifically, we calculated pairwise Tanimoto similarities within each superclass as defined in Table 2 of the reference, and derived a Representative Structural Similarity (RSS) score by averaging the top five similarity values for each compound.

[0070] This analysis showed that several superclasses — most notably the Benzenoids — contain compounds with relatively low structural similarity to other members of the same class. For example, 11-beta-Prostaglandin F2 alpha, classified as a Benzenoid in the referenced work, exhibits an RSS of approximately 0.5 within that group. Interestingly, it shows a much higher Tanimoto similarity (e.g., 0.9) to compounds in the “Lipids and lipid-like molecules” superclass, such as methyl linoleate.

[0071] These findings underscore a key limitation of semantic grouping: structurally dissimilar compounds may be grouped together, while structurally similar ones may be excluded. In contrast, our approach — based solely on molecular similarity — yields more homogeneous training sets, which in turn support more effective model training and improved prediction accuracy.

[0072] Benzenoids

[0073] LIST OF REFERENCE SIGNS

[0074] LC liquid chromatography

[0075] DDA data dependent acquisition DIA data independent acquisition

[0076] CCS collision cross-section

[0077] PASEF Parallel accumulation-serial fragmentation

[0078] RSS representative structure similarity TIMS trapped ion mobility spectrometry

[0079] SRM Selected Reaction Monitoring

[0080] MRM Multiple Reaction Monitoring

Claims

CLAIMS1. A computer implemented method of training machine learning model for the predictive calculation of a collision cross-section (CCS) value of a molecular structure belonging to a similarity structure family, comprising: collecting data from an existing database with at least molecular structure information and associated values of collision cross-section values, in that for training for a model for that similarity structure family a training subset from that database is generated for said similarity structure family, in that a) molecular fingerprints are calculated for the molecular structures in the database, and b) the calculated molecular fingerprints are clustered in similarity groups of fingerprint similarity, c) and for training the machine learning model for prediction of a similarity structure family only the information of molecular structures belonging to the same similarity group is used.

2. Method according to claim 1, wherein the similarity groups are similarity groups of highest fingerprint similarity, and preferably are formed around chemical substructure units.

3. A computer implemented method for the predictive calculation of a collision cross-section value of a specific molecular structure, wherein for the calculation of the collision cross-section of that specific molecular structure a) in a first step the molecular fingerprint of that specific molecular structure is calculated; b) in a second step the fingerprint similarity between the fingerprint of that specific molecular structure calculated in a) and at least one, preferably at least 10, or at least 100, or all structures of the training data for several similarity structure family prediction models is calculated; c) for the predictive calculation of the collision cross-section of that specific molecular structure the model for the similarity structure family having the highest fingerprint similarity established in step b) is used, or models of several similarity structure families with the highest fingerprint similarity established in step b) are used for calculation and used for prediction and the final predicted collision cross-section value is taken as an average,weighted average or median value of these calculated predictive values.

4. Method according to claim 3, wherein in step c) a model obtained in a method according to any of claims 1 or 2 is used.

5. Method according to any of the preceding claims, wherein the molecular fingerprint is a binary substructure fingerprint in the form of an ordered list of binary bits, wherein each bit represents a Boolean determination of or test for the presence of at least one of an element count, an element type, connection type, environment, in the chemical structure, wherein preferably the substructure fingerprint is according to the PubChem standard.

6. Method according to any of the preceding claims, wherein the fingerprint similarity in step b) is established using a, preferably binary, substructure fingerprint vector as basis for establishing the fingerprint similarity, and wherein fingerprints are compared by a numerical evaluation of the concordance of substructure fingerprint vectors of different structures.

7. Method according to any of the preceding claims, wherein the fingerprint similarity in step b) is established using the Tanimoto coefficient.

8. Method according to any of the preceding claims, wherein the fingerprint similarity in step b) is established using the highest Tanimoto coefficient between the specific molecular structure and the structures used to train the respective machine learning model.

9. Method according to any of the preceding claims, wherein the fingerprint similarity in step b) is established using an average of the two, three, four, five or ten best Tanimoto coefficients between the specific molecular structure and the structures used to train the respective machine learning model.

10. Method according to any of the preceding claims, wherein the fingerprint similarity in step b) of the method according to claim 1 or claim 3 is established using the Tanimoto coefficient, and wherein for clustering, structures having a similarity of at least 0.8, or at least 0.9 or at least 0.95 are used.

11. A method of data dependent or data independent combined liquid chromatography (LC), ion mobility and tandem mass spectroscopy analysis, comprising the following steps: introducing precursor ions resulting from LC separation into at least one trapped ion mobility spectrometry (TIMS) separator, and separating the precursor ions according to mobility in the trapped ion mobility spectrometry (TIMS) separator, sequentially releasing precursor ions from said trapped ion mobility spectrometry (TIMS) separator according to their ion mobility, introducing said released precursor ions into a mass filter which selectively transmits precursor ions having m / z values falling within a controllable mass window, fragmenting the precursor ions transmitted through said mass filter to generate fragment ions, carrying out a mass spectroscopy measurement on said fragment ions, wherein each fragment ion is associated with a mass window and an ion mobility (IM) range, and associating detected fragments with its corresponding precursor ion, wherein said trapped ion mobility spectrometry (TIMS) separator and said mass filter are scheduled in a synchronized manner such as to carry out a plurality of ion mobility (IM) scans, during which precursor ions of increasing or decreasing IM are successively released from said second ion mobility separator (IMS), and during which the mass window of said mass filter is shifted continuously or stepwisely towards lower or higher m / z values, respectively, wherein for annotation of signals, collision cross-section values of a specific molecular structure from databases not including collision cross-section information, crosssection value prediction for that specific molecular structure according any of the preceding claims 3-10 is used.

12. A method according to claim 11 , wherein at the beginning of one LC observation retention time window a survey scan is taken, wherein in that survey scan a full ion mobility width of interest and a full m / z width of interest is scanned, and wherein from that survey scan automatically a list of detectable signals is determined, from that list of detectable signals the strongest detectable signals are automatically selected for scheduling at least once in a synchronized manner in that LC observation retention time window.

13. Method according to any of the preceding claims 11 or 12, wherein said stepof associating a detected fragment with its corresponding precursor ion is based on determining or utilizing the corresponding mass windows and ion mobility (IM) ranges associated with various occurrences of said fragment in said mass spectrometry measurement.

14. Method according to any of the preceding claims 11-13, wherein in said IM scans, adjacent mass windows that are associated with consecutive mass spectroscopy measurements of fragment ions overlap, such that the precursor ions transmitted through said mass filter during one IM scan are located in at least one continuous scan region in an m / z-IM plane which extends in a generally diagonal direction in said m / z-IM plane, wherein adjacent scan regions associated with different IM scans overlap in the m / z-direction.

15. Use of a method of any of the preceding claims for the analysis of a complex sample, including samples in proteomics and in particular metabolomics.

Citation Information

Patent Citations

  • Precursor selection for data-dependent tandem mass spectrometry

    US20190371585A1

  • Method and apparatus for data independent combined ion mobility and mass spectroscopy analysis

    US20220034840A1

  • Apparatus and method for parallel flow ion mobility spectrometry combined with mass spectrometry

    US7838826B1

  • High duty cycle trapping ion mobility spectrometer

    US9304106B1

  • Trapping ion mobility spectrometer with parallel accumulation

    US9683964B2