Structural elucidation from liquid chromatography tandem mass spectrometry data

The described workflow addresses the limitations of current LC-MS/MS methods by employing denoising and deconvolution techniques, coupled with consensus scoring and GPU-accelerated processing, to achieve high accuracy in structural elucidation of complex mixtures.

WO2026075614A1PCT designated stage Publication Date: 2026-04-09AGENCY FOR SCI TECH & RES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Current LC-MS/MS processing methods struggle with deconvoluting and denoising individual compounds in complex mixtures, leading to truncated analyses and incomplete structural elucidation, with up to 80% of signals remaining uninterpreted.

Method used

A workflow and system for automatically translating raw LC-MS/MS data into molecular structures using denoising, deconvolution, and consensus scoring, leveraging a library of tandem mass spectra and employing GPU-accelerated processing to identify molecular structures with high accuracy.

Benefits of technology

Achieves 81% Top 10 accuracy in predicting the correct molecular structure, enabling deeper exploration of complex mixtures at scale and accelerating discovery processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050543_09042026_PF_FP_ABST
    Figure SG2025050543_09042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method, and a system that executes such a method, for elucidation of molecular structures from raw liquid chromatography tandem mass spectrometry (LC-MS / MS) data. The method involves obtaining mass spectrum data, comprising MS1 and MS2 data, from the raw LC-MS / MS data. Compound peaks are then detected by denoising and deconvoluting individual compound fingerprints from the MS1 data. A type of charged adduct corresponding to each precursor molecular ion represented in the MS1 data is determined, and MS1 intensities of an isotope distribution signature for each charged adduct are then obtained. MS2 spectra can subsequently be extracted for each compound corresponding to a compound peak, based on each detected compound peak, charged adduct and MS1 intensities. From this information, a molecular structure can be elucidated from the MS2 spectra by comparing the MS2 spectra to a reference library of reference MS2 spectra and corresponding molecular structures.
Need to check novelty before this filing date? Find Prior Art

Description

STRUCTURAL ELUCIDATION FROM LIQUID CHROMATOGRAPHY TANDEM MASS SPECTROMETRY DATATECHNICAL FIELD

[0001] Embodiments of the present specification relate to the technical field of analyzing liquid chromatography tandem mass spectrometry (LC-MS / MS) data to elucidate the molecular structure of compounds.BACKGROUND

[0002] The process of isolating and identifying molecular structures from complex mixtures is a common pain point for high throughput LC-MS / MS analyses (e.g. in natural product discovery, untargeted metabolomics).

[0003] LC-MS / MS is a powerful technique for analysing complex mixtures in clinical diagnostics, food safety, forensic analysis, and many other applications. Due to its exceptional sensitivity, LC-MS / MS can often detect hundreds to thousands of compounds in a single run. These huge data volumes, coupled with limitations of current processing methods, often result in truncated analyses leaving up to ~80% of signals in complex mixtures uninterpreted.

[0004] Current LC-MS / MS processing methods encounter many roadblocks to success, including difficulty in deconvoluting and denoising individual compounds, incomplete MS / MS reference libraries limiting structural elucidation, and tedious manual intervention, limiting throughput and scale of in-depth analyses.|0005] It would be desirable, therefore, to provide a technology that circumvents this problem by directly identifying molecular structures from raw analytical data, thereby accelerating discovery, and enabling deeper exploration of complex mixtures at scale.SUMMARY|0006| Disclosed herein are workflows for automatically translating LC-MS / MS data, particularly raw LC-MS / MS data, of complex mixtures into their constituent molecular structures. Embodiments achieve Top 10 accuracy of 81% - i.e., these workflows are able to predict the correct molecular structure within the top 10 candidates >8 out of 10 times.

[0007] This technology leverages the library of tandem mass spectra and their corresponding molecular structures described in Singapore patent application No. 10202400652R, the entire contents of which is incorporated herein, by reference.

[0008] Disclosed herein is a method for elucidation of molecular structures from raw LC-MS / MS data of a sample, comprising: obtaining mass spectrum data, comprising data from the first stage of mass analysis (MSI) and data from the second stage of mass analysis in tandem mass spectrometry (MS2), from the raw LC-MS / MS data; detecting compound peaks by denoising and deconvoluting individual compound fingerprints from the MSI data; predicting a type of charged adduct corresponding to each compound peak represented in the MS 1 data; determining a monoisotopic mass from the predicted charged adduct type; obtaining MSI intensities of an isotope distribution signature for each charged adduct; extracting MS2 spectra for each compound corresponding to a compound peak, based on each detected compound peak, charged adduct and MS 1 intensities; and elucidating a molecular structure by consensus scoring based on at least two of a spectrum score derived by comparing the MS2 spectra to a reference library of reference MS2 spectra and corresponding molecular structures, an isotope score derived by comparing the isotope distribution signature to a reference library of calculated isotope distribution signatures of molecular formulae, and a monoisotopic mass comparison score derived by comparing the monoisotopic mass of the compound against a reference library of calculated monoisotopic mass of molecular formulae, and extracting the molecular structure from a reference library based on the consensus score.

[0009] Also disclosed herein is a system for elucidation of molecular structures from raw LC-MS / MS data of a sample, comprising: a receiver for receiving the raw LC-MS / MS data, the raw LC-MS / MS data comprising MSI and MS2 data; a peak detector for detecting compound peaks by denoising and deconvoluting individual compound fingerprints from the MS1 data; an adduct predictor for predicting a type of charged adduct corresponding to each precursor molecular ion represented in the MSI data, and for determining a monoisotopic mass from the predicted charged adduct type; an isotope analyser for obtaining MSI intensities of an isotope distribution signature for each charged adduct;a MS2 spectrum extractor for extracting MS2 spectra for each compound corresponding to a compound peak, based on each detected compound peak, charged adduct and MSI intensities; and a molecular analyser for elucidating a molecular structure by consensus scoring based on at least two of a spectrum score derived by comparing the MS2 spectra to a reference library of reference MS2 spectra and corresponding molecular structures, an isotope score derived by comparing the isotope distribution signature to a reference library of calculated isotope distribution signatures of molecular formulae, and a monoisotopic mass comparison score derived by comparing the monoisotopic mass of the compound against a reference library of calculated monoisotopic mass of molecular formulae, and extracting the molecular structure from a reference library based on the consensus score.

[0010] Consensus scoring can be 3-part consensus scoring using the spectrum score, isotope score and monoisotopic mass comparison score.

[0011] Embodiments of the workflow herein employ a peak detection algorithm and perform deconvolution from MS I spectra of mixtures. These processes provide a novel way of deconvolution and denoising independent of peak resolution, that does not require resolved peaks, employing examination of individual m / z’s over time to characterise compound signals.

[0012] Embodiments of the methods, and systems implementing those methods, described herein provide 3-part consensus scoring for structural elucidation - i.e., determining molecular structure from tandem mass spectrometry data. This involves a combination of 3 key features (MS2, isotopic distribution, monoisotopic mass) weighted by a tuned set of parameters for optimal accuracy. The inclusion of multiple characteristic features improves accuracy of structural elucidation.

[0013] Embodiments of the methods, and systems implementing those methods, described herein enable GPU-accelerated search algorithms. Spectra similarities can be computed in parallel using CUDA tensor operations. Highly parallelized CUDA tensor operations can be utilized to match query spectra to shortlisted library spectra.BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Some embodiments of the present invention will now be described, by way of nonlimiting example only, with reference to the accompanying drawings in which:

[0015] FIG. 1 is an overview of an automated Workflow for Intelligent Structural Elucidation (WISE), used to translate raw LC-MS / MS data of a mixture into its constituent molecular structures.

[0016] FIG. 2 depicts graphs of normalized MS I intensity of selected m / z’s (mass-to- charge ratio) against retention time (min) for three different types of signals - “noise-like”, “fragment-like”, and “compound-like”.

[0017] FIG. 3 shows: (A) example mass spectrum (MS 1) being evaluated with two example adduct hypotheses of [M+H]+(Guess #1) and [M+Na]“ (Guess #2); (B) example of detailed spectral comparison for adduct hypotheses and explained adduct scores based on those hypotheses.

[0018] FIG. 4 provides weighted cosine similarity comparisons of the tandem mass spectrometry (MS2) spectra of: (A) ethyl acetoxy(diethoxyphosphoryl) acetate at 99% similarity, and, (B) 2,2’ -dipicolylamine at 67% similarity, to the predicted spectra in the spectral reference library.

[0019] FIG. 5 is a ternary plot showing how rank of correct molecular structure varies with different combinations of weights for the components of MS2 similarity, isotope distribution, and monoisotopic mass

[0020] FIG. 6 shows the Top 3 candidates predicted from structural elucidation of LC- MS / MS data of Staurosporine.DETAILED DESCRIPTION

[0021] Described is a fully automated workflow that can leverage any spectral reference library and GPU-accelerated processing to translate raw LC-MS / MS data of complex mixtures into their constituent molecular structures. Some embodiments achieve 81% Top 10 accuracy with high throughput and at scale.

[0022] The broader workflow 100 for elucidation of molecular structures from raw LC- MS / MS is set out in FIG. 1. The workflow 100 comprises a data pre-processing stage 102, a spectral characterization and data extraction stage 104, and a GPU-accelerated MS2 library search for structural elucidation 106, to output a predicted molecular structure 108.

[0023] The data pre-processing 102 stage obtains mass spectrum data, comprising MSI and MS2 data, from the raw LC-MS / MS data. LC-MS / MS equipment typically produces data in their own vendor-specific proprietary format (e.g. d for Agilent, raw for ThermoFisher). The present workflow is vendor-agnostic, facilitating cross-vendor compatibility to enable workflow 100 to take in data from various vendors. Particularly by conversion to an open- source data format and using a reference library where mass spectra are predicted for various collision energies (eV), LC-MS / MS data can be matched to a real or predicted spectrum regardless of the collision energy or equipment on which the data was generated.

[0024] At step 110, MS I data is processed by extracting MSI spectra for a plurality of timestamps (retention time or sampling time) from the raw LC-MS / MS data. Tn some embodiments, the MSI spectra are extracted for every timestamp. This can be done using any appropriate software, such as MSConvert. The resulting MS 1 spectra and timestamps form an output that is provided as an input for the next step of the structural elucidation workflow.

[0025] At step 112, LC-MS / MS analytical data (raw or otherwise) is processed to extract MS2 data. The MS2 data can be extracted to any desired format, such as Mascot Generic Format (.mgf) files using MSConvert. The output of step 112 comprises a plurality of, or list of, MS2 spectra with associated timestamps. That output is provided as input for the next step of the structural elucidation workflow.

[0026] Stage 104 involves spectral characterisation and data extraction from the outputs of steps 110 and 112, namely the MSI and MS2 spectra with associated timestamps. These characterization and extraction steps use a series of data analysis algorithms that deconvolute individual compound fingerprints and extract spectral features. Tn particular, compound peaks are detected by denoising and deconvoluting individual compound fingerprints from the MSI data. The deconvoluted list of compound fingerprints, or peaks, are then passed on for structural elucidation in the final step of the workflow.

[0027] Denoising and deconvoluting processes make use of a denoising detector 114. The denoising detector 114 looks at individual m / z’s and evaluates characteristics thereof. Current commercial LC-MS / MS processing applications typically deconvolute (i.e., determine which features of a mass spectrum belong to which precursors) via peak picking of total ion current (TIC) chromatograms representing the summed intensity across the entire range of m / z’s detected. This method is limited by the resolution between different peaks (as some may coelute), and the signal -to-noise ratio (as summed background noise can at times obscure weak sample signals). To address these issues, the denoising detector denoises and deconvolutes individual compound fingerprints by looking at individual m / z’s and evaluating their characteristics (FIG. 2).

[0028] The denoising detector 1 14 may generate a graph of MSI intensities of individual m / z’s against retention time. This graph may be normalized. Normalization may address varying instrument sensitivities, sample concentrations, and injection parameters. Varying instrument sensitivities, sample concentrations, and injection parameters may give rise to MS intensities varying from 10E5 to 10E8 on a typical spectrum. Normalization can therefore enable use of a percentage (of peak intensity) rather than absolute values, the latter being subj ect to variation. In some embodiments, normalization is against a reference spectrum. In otherembodiments, normalization is against its own spectrum - e.g., spectral band of spectrum or fixed AUC. Characteristics of the graph are ascertained to categorise the graph as representative of a compound or noise. This can involve, for each m / z, categorising it as belonging to a noiselike category, a compound-like category, or a fragment-like category. These three categories are defined as:• “Noise-like”: signals with similar intensities across the entire time range, indicating a noise or background signal. The present framework focusses on detecting compound-like signals But, distinguishing compound-like signals from other signals requires an understanding of those other signals - i.e., what is not a compound-like signal.• “Fragment-like”: signals with more than 1 major peak, usually localized around the same retention time. This indicates a common scaffold or fragment from compound analogues eluting at different distinct times. Fragment-like signals may be detected by identifying one or more peaks aside from the detected peak that have a summed MSI intensity of at least 50% of the summed MSI intensity of the detected peak at the same m / z. This includes peaks that are separated from each other along the spectrum.• “Compound-like”: signals following a regular peak shape distribution, in-line with compound elution from the column. No other significant signals for this mass are observed elsewhere on the spectrum, indicating a unique molecular ion.

[0029] Denoising may further comprise discarding graph categorized as anything other than the compound-like category. The denoising step therefore eliminates signals from fragments as well as background noise signals, effectively denoising the spectrum. Fragment-like signals indicate that the MSI precursor m / z is not that which is applicable for a whole molecule. Instead, it is of a fragment In this context, the whole molecule precursor m / z is the relevant information, as it can then be analyzed to elucidate molecular structure. Analysis of fragments may provide information on compounds present, but are not as relevant for structural elucidation.

[0030] Deconvoluting individual compound fingerprints is a process whereby the mass spectra for different compounds are distinguished from each other. To deconvolute, for each m / z the “compound-like” peak, represented by the graph for a particular m / z retained after denoising, are filtered by considering the area under the curve of summed intensities (usually normalized intensities) in a region around, and comprising, the peak (i.e., the highest peak on the graph) versus the area under curve of the rest of the spectra for a specific m / z. In particular,deconvoluting individual compound fingerprints may comprise, for each “compound-like” peak of a unique mass, characterizing the peak according to:1. Area under the curve of a region comprising the peak - the normalized area must be less than a certain threshold to indicate a sharp, well-defined peak shape (the threshold may be pre-set at 3, where larger values mean broader peaks are recognized as compounds. If the threshold is too large, then areas containing multiple peaks will be recognized as a single peak). This is determined by increments in retention time before and after the highest peak. This is currently set at 10 timesteps from the highest peak intensity. Meaning the algorithm will search from the timestamp (residence time) of the highest peak intensity to +10 timesteps away. This parameter is also adjustable. The timestep is relative to the instrument acquiring the data, this can range from 0.01 ~ 0.02 min (as observed so far).2. Area under the curve outside the peak region - the normalized area must be less than a certain threshold to indicate a significant unique signal (the threshold may be pre-set at 0.5, this means that there are minimal peaks outside of the peak region)

[0031] These additional filters allow for fine-tuning to recognize “compound-like” signals that may not fit with standard peak shapes (e.g. more broad peaks). Particularly, peaks with a normalized area under curve above a pre-set threshold for a region comprising the peak may not be recognized as compounds, and peaks with a normalized area under curve above a preset threshold for curves outside the peak region may be attributed to noise or fragment-like signals. Moreover, evaluating individual m / z’s this way allows for a deeper deconvolution as it enables overlapping “compound-like” signals to be separated and low intensity “compoundlike” signals to be differentiated from background noise.

[0032] The present methodology performs a deep deconvolution at the individual m / z level. This enables separation of overlapping “compound-like” signals based on signal characteristics, such as calculating AUC and using retention time thresholds (within a predetermined retention time of the peak) to determine if the peak has a shape indicating presence of a compound, and thereby distinguishing it from other peaks representing other compounds. This process is more reliable for complex mixtures, when compared with relying on peak picking of total ionchromatogram (TIC) spectra that encounters difficulties when faced with overlapping compound spectra.

[0033] Current commercial LC-MS / MS processing applications typically deconvolute (i.e., determine which features of a mass spectrum belong to which precursors) via peak picking of TIC chromatograms representing the summed intensity across the entire range of m / z’s detected. Overlapping peaks thus occur when there are two or more compounds that elute around the same retention time. As the TIC chromatogram is the summed intensity across all m / z’s, the peaks overlap and are hard to resolve.

[0034] Using, per the present methodology, intensity over time at the m / z level by creating a chromatogram for each m / z, and analysing peaks per m / z chromatogram, two (or more compounds) that elute around the same retention time can still be resolved, assuming they are of different m / z’s. To illustrate, given two hypothetical compounds:Compound A - m / z 619.2531, retention time 3.25 minCompound B - m / z 552.3162, retention time 3.25 min

[0035] The TIC spectra will show a single peak at 3.25 min. In contrast, the present workflow will exhibit two peaks: one for m / z 619.2531 at 3.25 min; and one for m / z 552.3162 at 3.25 min. In this way, the two compounds can be resolved.

[0036] This method allows for comparison against reference libraries to predict a compound structure.

[0037] After denoising and deconvolution, the fingerprinted compounds are precursor molecular ions, outputted from the denoising detector 114. The adduct identifier 116 is then used to predict a type of charged adduct corresponding to each precursor molecular ion represented in the MSI data and thus reflected in the output from the denoising detector 114. Understanding the adduct type of parent molecular ions observed is critical for subsequent structural elucidation as it determines compound monoisotopic mass. Adduct scoring may employ any desired method, provided the method selected is then consistent for all analyses. For example, a heuristic scoring method may be used for interpreting electrospray ionization (ESI) mass spectra, by which MS spectra are scored based on intensity, mass accuracy, and isotope charge agreement of adducts and related ionization products, the heuristic scoring method annotating peaks of the (de)protonated molecule and adduct ions - see., e.g., the findMAIN algorithm.

[0038] The Adduct Identifier 116 uses a MSI adduct scoring algorithm that evaluatesadduct identity hypotheses based on total explainable MS I peaks (FIG. 3) - i.e., peaks that are expected to be present if the adduct is correct The algorithm is run for a pre-defined list of adducts, and an adduct score is produced for each adduct. The highest scoring adduct may then be taken as the predicted adduct type. The predicted adduct type may then be used to calculate a monoisotopic mass of the compound from the MSI precursor m / z. In some embodiments, charged adducts are predicted based on a summed intensity of all explained adducts relative to a deisotoped MSI spectrum derived from the MSI data, a calculated adduct score, and the number of explained adducts for a current hypothesis, of one or more hypotheses, relative to a maximum number of explained adducts detected in any of the hypotheses. Deisotoped MSI spectra are obtained by taking a MSI spectrum, analysing raw MSI data to extract isotopic patterns of precursor ions and removing components of the MS 1 spectrum contributed to by those isotopic patterns. In this regard, explained adduct scores are calculated by the following equation:Score = 50% * Intensity + 10% * Isocharge + 40% * AdductsWhere, the Intensity refers to the modified fraction of the summed intensity of explained peaks (peaks to which an adduct can be assigned, from a reference dictionary of adducts), over the summed intensity of all the peaks in the MSI spectra, Isocharge refers to the ratio of the sum of all the charge validity coefficients for all the explained peaks (where “1” indicates the observed charge matched the predicted charge based on the annotated adduct, “0.75” indicates the observed charge is 0, and “0” indicates all other cases), over the total number of MS 1 peaks, and Adducts refers to the fraction of MSI peaks which are explained (by being able to assign a plausible adduct from a reference dictionary of adducts based on the adduct hypothesis of the main MSI peak being analysed) over all the peaks in the MS1 being analysed - “explainable adducts”.

[0039] The weight percentages in the above formula were empirically derived from >1,000 test spectra, provided in literature. The skilled person will appreciate that the weights may vary. For example, the weight for Intensity may be between 35% and 65%, preferably between 40% and 60%, and more preferably between 45% and 55%, the weight for Isocharge may be between 5% and 15%, preferably between 7% and 13% and more preferably between 9% and 11%, and the weight for Adducts may be between 35% and 45%, preferably between 37% and 43% and more preferably between 38% and 42%.

[0040] The Isotope Characterizer 1 18 leverages the predicted charged adducts from theAdduct Identifier 116, to obtain MS I intensities of an isotope distribution signature for each charged adduct. The isotope distribution signature may be obtained based on the identified precursor m / z (obtained from the earlier denoising deconvolution module 114) at the MSI of a specific timestamp, where the specific precursor m / z is located. Successive M+l, M+2, M+3, etc., peaks may be identified based on peaks which are +1 (corresponding to a charge of 1), +0.5 (corresponding to a charge of 2), +0.33 (corresponding to a charge of 3), and so on, from the main peak. The natural isotopic abundance of elements translates to a unique isotopic distribution signature (pattern of peaks in MSI Spectra) for each molecular formula. The isotope distribution signature is used to differentiate different molecular formulae between compounds of similar monoisotopic masses. Based off the proposed charged adduct predicted from the Adduct Identifier 116, MSI intensities of the isotope distribution signature are extracted for the detected compound to be passed on as input for the next step.

[0041] Sequentially applying the denoising detector 114, adduct identifier 116, and isotope characterizer 1 18 modules yields a deconvoluted compound list of monoisotopic ion masses and their respective retention times. The retention time is determined via the denoising detector 1 14 module. For a given m / z, after normalization, the timestamp where the highest intensity occurs (i.e. value 1) is taken as the retention time. Based on the compound peaks (“compoundlike” signals retained after filtering by the Denoising Detector 114), charged adducts identified by the Charged Adduct identifier 118, and MSI intensities extracted by the Isotope Characterizer 118, MS2 spectra are extracted by the Spectral Extractor 120, for each compound corresponding to a compound peak.

[0042] The MS2 Spectral Extractor 120 reads the retention time and precursor m / z of the detected compound, scans through previously extracted MS2 data, and extracts a corresponding MS2 spectra generated from the fragmentation of the MSI precursor m / z of the detected compound. The scanning and extraction processes will be known to the skilled person, in view of present teachings. These characteristic MS2 spectra are used for structural elucidation in the final step of the workflow.

[0043] Thus, at the conclusion of the spectral characterisation and data extraction stage 104, a deconvoluted list of compound fingerprints with the following key information is obtained for subsequent structural elucidation:1. Retention Time (in minutes, min) - provides information on compound chemical profile based on affinity to the column used for liquid chromatography.2. Monoisotopic mass (in m / z) - observed exact monoisotopic mass observed for parent molecular ion, provides information on molecular weight and molecular formula.3. Adduct type - predicted adduct type, provides information on MS2 fragmentation fingerprint, molecular weight, and molecular formula.4. Isotopic distribution - observed MSI intensities for the isotopic distribution of the molecular ion, provides information on molecular formula to differentiate between similar monoisotopic masses.5. MS2 spectra - fragmentation fingerprint of the molecular ion, provides information on specific molecular structure.

[0044] The above outputs are passed to the structural elucidation stage 106. In the structural elucidation stage 106, the molecular structure are elucidated from theMS2 spectra by consensus scoring, comparing the MS2 spectra to a reference library of reference MS2 spectra and corresponding molecular structures, comparing isotope distribution signatures to a reference library of calculated isotope distribution signatures of molecular formulae, and comparing monoisotopic mass of the compound against a reference library of calculated monoisotopic mass of molecular formulae. Consensus scoring will generally be 3-part consensus scoring - i.e., using all three of the spectrum score, monoisotopic mass comparison score and isotope distribution score. This MS2 library search can be GPU-accelerated

[0045] Structural elucidation from MS2 spectra typically involves comparison against spectral reference libraries for the closest match. Leveraging the reference library described in Singapore patent application No. 10202400652R, or other reference libraries, a match can be obtained with high accuracy. In addition, the present methodology incorporates additional spectral comparison metrics (i.e. isotopic distribution and monoisotopic mass) to increase the accuracy of structural elucidation.

[0046] Subsequent fragmentation of MSI molecular ions in tandem mass spectrometry (MS2) provides signature fingerprint spectra related to molecular structure. By matching reference spectra to these fingerprint spectra, the exact molecular structure of the sample can be elucidated. This matching process is performed using MS2 Similarity module 122 of a molecular analyser that also comprises an Isotope Distribution module 124 and Monoisotopic Mass module 126. The choice of similarity algorithm used by the MS2 Similarity module 122 is critical as it will be used to separate the correct spectra from similar ones in a spectral reference library, where those spectra differ based on the equipment used, collision energy, and fragments generated during collision. Any appropriate similarity metric may be employed. For present purposes, a weighted cosine similarity algorithm is used, placing greater importance on heavier m / z signals according to the following equation:Weighted Cosine Similarity(lq, Ij)where Tqand h re vectors of m / z intensities representing two spectra (the spectrum from the reference library and the spectrum being compared thereto), mk and Ik are the m / z and intensity found at m / z = k, Mqand Mi are the largest binned values (e g , rounded to nearest integer) of Iqand L with non-zero values, and Mmax is the larger of Mqand Mi. Other similarity metrics may be employed, including Euclidean distance, Absolute Value Distance, unweighted cosine similarity, dot product, etc.

[0047] In some embodiments, the cosine similarity is weighted by m / z. This is because the heavier peaks in mass spectra corresponding to larger fragments are more characteristic and useful in practice for structural elucidation. Each intensity may be weighted by multiplying itself with its associated m / z value, hence linearly weighting each peak intensity by its m / z.

[0048] FIG. 4 illustrates results of weighted cosine similarity comparisons of the tandem mass spectrometry (MS2) spectra of: (A) ethyl acetoxy(diethoxyphosphoryl) acetate at 99% similarity, and; (B) 2,2’ -dipicolylamine at 67% similarity, to the predicted spectra in the spectral reference library.

[0049] As a benchmarking study, 3 in-house experimental MS2 samples of known compound standards were searched against a reference MS2 library using both non-weighted and weighted cosine similarity algorithms. The results are set out in Table 1.Table 1. Summary of results from searching generated reference MS2 library (Singapore patent application No. 10202400652R) with weighted and non-weighted cosine similarity algorithms.*rank = MS 2 reference library is filtered by ±10 m / z around the precursor m / z then ranked in descending order by (weighted) cosine similarity score, rank of the correct molecular structure is then reported.

[0050] Generally, weighted cosine similarities give higher similarity values for the correct structure as well as a correspondingly higher rank when compared against other spectra in the reference library, increasing overall accuracy in MS2-to-structure elucidation. These weighted cosine similarity scores are brought forward as one of the metrics for overall determination of structure from MS2.

[0051] The Isotope Distribution module 124 determines a similarity between the observed MSI isotope distribution, and calculated MSI isotope distributions of molecules in the reference library. The Isotope Distribution module 124 may apply a standard cosine similarity formula to calculate the similarity between the observed MSI isotope distribution and the calculated MSI isotope distributions of the molecules in the reference library. These isotope distribution similarity scores (cosine similarities) provide an indication of the molecular formula of the sample and further constrain the possible molecular structures it may take. These isotope distribution scores are then brought forward as one of the metrics for overall determination of structure from MS2.

[0052] To further refine the structural elucidation, the Monoisotopic Mass module 126 constrains the possible molecular structures the compound from the sample can take. High resolution mass spectrometry can provide extremely accurate (typically up to 4 decimal places) monoisotopic masses for the molecular ions being examined. With knowledge of the adduct type predicted by the Adduct Identifier 116, the exact monoisotopic mass of the molecular ion is a powerful source of information to constrain the possible molecular structures it could refer to. A monoisotopic mass score, also referred to as a monoisotopic mass comparison score, is calculated by the Monoisotopic Mass module 126 according to the following formula:Monoisotopic mass scorewhere, the “mass filter” is an m / z threshold around the molecular ion by which the reference library is filtered, the “observed mass” is the exact monoisotopic mass of the molecular ion observed, the “reference mass” is the exact monoisotopic mass of the molecular ion in the reference library and “abs” refers to the figure being absolute or in ‘magnitude’ form without sign. The m / z threshold may be empirically determined. For present purposes, the empirically determined m / z threshold is 0.005 m / z, wherein a lower value would not account for possibleexperimental errors, and a value which is set too high may result in more reference molecules being evaluated or may increase computational costs and accuracy.

[0053] The monoisotopic mass comparison scores are then calculated for all potential candidate molecular structures within the selected threshold around (i.e., within the mass filter) the observed monoisotopic mass of the sample molecular ion and brought forward as one of the metrics for overall determination of structure from MS2.

[0054] The molecular analyser performs consensus scoring using the 3 scores (MS2 similarity, isotope distribution score, monoisotopic mass comparison score) calculated for each potential candidate molecular structure in the reference library (i.e., calculated monoisotopic mass, calculated isotope distribution, and predicted MS2 spectra). In some embodiments, these scores are evaluated into a consensus score by weighting each score, and the highest scoring entry is taken as the de facto predicted structure. Using these scores, the molecular structure can be extracted from the reference library.

[0055] The optimized weights were fine-tuned against 18 in-house experimental MS2 spectra of known compounds (Table 3 - all entries excluding fragmented molecules) to maximise accuracy. Weights across the 3 components were varied at 1% increments to give 5,151 possible combinations. These combinations were applied, and the average of the correct molecule ranks were taken as an indication of performance for that weight combination (FIG. 5). FIG. 5 shows how varying the weights across the 3 components (MS2, isotope, mass) affects the rank of the correct structure, wherein the rank is obtained from the 3 -part consensus scoring of the experimental data against the reference library as discussed above. 695 combinations gave identical top performance of average rank = 4.39, indicating a range of weight values for optimal performance.

[0056] A statistical analysis of these combinations shows:Average optimal MS2 similarity weight: 52.3% ± 10.1% Average optimal isotope distribution weight: 16.8% ± 11.4% Average optimal monoisotopic mass weight: 30.9% ± 6.2%

[0057] For comparison, the performance of relying on a single score is as follows:Average rank with MS2 similarity only: 4.78Average rank with isotope distribution only: 20.3Average rank with monoisotopic mass only: 11.6Average rank with optimized consensus score: 4.39

[0058] In-line with theory, MS2 similarity carries the most information about the exact molecular structure and thus should hold the highest weightage. Both isotope distribution and monoisotopic mass are supporting constraints to help focus the search space to the most likely candidates. However, due to the presence of noise signals that could potentially disrupt observation of an ideal isotope distribution, a slightly higher weightage is given to the monoisotopic mass comparison score. Overall, experimentally optimized consensus scores are calculated according to the following formula:Consensus score= 50% * MS2 Similarity + 20% * Isotope Distribution + 30%* Monoisotopic Mass

[0059] To accelerate the processing time, code was developed to leverage highly parallelized computing for MS2 similarity calculations, to compare thousands of MS2 signals simultaneously 10 timing runs were done to benchmark processing time to compare a sample MS2 spectra against 100 reference MS2 spectra.For 100 MS2CPU version = 209 ± 1.3 msGPU version = 3.66 ± 0.42 msFor 68 million MS2 (size of in-house reference library)CPU version = 2,368 min (1.64 days)GPU version = 41.5 minThe GPU version showed a significant 57-fold acceleration in processing time over the CPU version, shortening the wait time from days to minutes when performing spectral comparisons at scale.

[0060] Experimental validation

[0061] To validate the accuracy of the workflow of FIG. 1 in performing structural elucidation two different benchmarking studies were conducted. The first study was run on aset of 21 in-house validated LC-MS / MS data of known compounds, and the second study was run on publicly available MS2 data retrieved from the Global Natural Products Social Molecular Networking (GNPS) database.

[0062] The results of the first study, namely running the workflow of FIG. 1 on the 21 samples are summarized below in Table 3.Table 3. Summary of results from running the automated workflow for intelligence structural elucidation (WISE) on 21 in-house LC-MS / MS samples of known compounds.Adduct Type = predicted adduct type. Fragmented = no molecular ion found. WISE Rank = the rank of the correct structure in the list of predicted candidate molecular structures. WISE Score - consensus score of the correct molecular structure.

[0063] Out of a total of 21 samples, the present workflow achieved 52% Top 1 accuracy (i.e. the correct structure is selected as the top candidate). However, when not Top 1, the present workflow still predicts a structural isomer possessing the same scaffold (FIG. 6).

[0064] Of the 21 samples, 33% (7 out of 21) fall in this category where structural isomer analogues occupy the Top 1 position. As these structural isomers also significant structural information, this produces a total of 85% Top 1 accuracy for significantly informative molecular structures. The remaining 15% (3 out of 21) failed as no molecular ion could be detected. If a hard ionisation method is used or when an unstable compound is ionised, the resulting ion can fragment into smaller, more stable ions. In the event where no significant amount of the parent structure is detected, it will not be picked up for MS2 analysis and hence no information on the full molecular structure is obtained by LC-MS / MS analysis - rendering structural elucidation impossible.

[0065] For a more rigorous evaluation, the present workflow was applied to three publicly available natural product libraries retrieved from the Global Natural Products Social Molecular Networking (GNPS) database, being the GNPS Community Library - 1,327 MS2 spectra fromuser contributions, the NIH Natural Products Library Round 1 - 589 MS2 spectra from the National Institutes of Health (NIH) Natural Products Library, and the NIH Natural Products Library Round 2 - 2,589 MS2 spectra from the National Institutes of Health (NIH) Natural Products Library.

[0066] Due to the scarcity of publicly available data however, this broader validation study had the following limitations:1. As no MSI data was publicly reported, only MS2 data was obtained for each compound. This prevents implementation of the “Isotope Distribution” metric, reducing the consensus score to just 2 metrics of “MS2 Similarity” and “Monoisotopic Mass” which reduces accuracy.2. Collision energy of the respective MS2 spectra were not reported. Collision energy determines the resulting MS2 fragmentation pattern and has critical importance on how the “MS2 Similarity” metric is calculated. In the absence of any information, we assumed that collision energies applied followed a typical “normalised collision energy” ramped method where higher collision energies are applied to larger masses to maintain a reasonable amount of fragmentation. However, incorrectly assumed collision energies would reduce accuracy.3. The MS2 data retrieved were filtered to only include entries which were also present in our spectral reference library. This was done to only include molecules which could be ranked, as any not present within the reference library would end up with no matches.Table 4. Summary of results from running the automated workflow for intelligence structural elucidation (WISE) on MS2 spectra from three public libraries retrieved from the Global Natural Products Social Molecular Networking (GNPS) database.

[0067] The present workflow achieved overall 81% Top 10 accuracy (i.e. the correctmolecular structure was ranked within the top 10 candidates), with an average rank of 11 ± 30.5, and a median rank of 2 (Table 4). The high standard deviation and low median rank indicate high fluctuations for selected samples as seen with the previous validation study. A deeper analysis showed that this is linked to the mass of the sample, with higher accuracies for samples with higher molecular weight (Table 5).Table 5. In-depth analysis of results based on sample molecular weight.m / z = mass to charge ratio.

[0068] For molecules heavier than 500 m / z, the Top 10 accuracy of the present workflow further increases to 92%. This effect can be attributed to the diversity of structures present at lower molecular weights, where there are many more known, characterized isomers to dilute the rank This effect has also been observed for the in-house data validation study.

[0069] The present workflow therefore provides the following key technical features for intelligent structural elucidation from MS2 spectra of complex mixtures:1. A pre-processing step to accept vendor-agnostic data input. For example, MSConvert and other software may be used to transform vendor-specific data formats into an open- source format like mgf so the workflow can run on multiple formats.2. A peak detection algorithm for denoising and deconvoluting individual compound fingerprints from complex MSI spectra of mixtures3. A weighted cosine similarity algorithm that places more importance on higher molecular weight signals for MS2 spectra4. An isotope characterizing algorithm to identify isotope distributions5. An adduct identifier algorithm to predict the adduct type of molecular ions6. 3-part consensus scoring and formula (MS2 similarity, isotope distribution, monoisotopic mass) to rank candidate structures7. A GPU-accelerated processing algorithm for MS2 reference library search

[0070] This technology allows for the direct interpretation of raw LC-MS / MS data of complex mixtures to its constituent compound molecular structures which can significantlyaccelerate efforts in the following areas:• Identification of novel natural products for high-value specialty chemicals, for example, antimicrobials, photoactive molecules, mosquito repellents• Accelerate natural product discovery pipelines• Elucidation of new metabolites as high-confidence biomarkers for various diseases, for e g. pancreatic cancer• Determination of food safety by labelling metabolomics data

[0071] For the sake of convenience and illustration in description, the aforementioned method embodiments are all expressed as a combination of a series of actions, but those skilled in the art should be aware that the embodiments of the present specification are not limited to the described order of actions, because according to the embodiments of the present specification, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the present specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present specification.

[0072] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0073] The preferred embodiments of the present specification disclosed above are only used to help illustrate the present specification. The optional embodiments do not describe all details exhaustively nor limit the invention to the specific embodiments described. Obviously, many modifications and changes may be made based on the content of the embodiments of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and use the present specification. The present specification is limited only by the claims appended hereto along with their full scope and equivalents.

Claims

CLAIMS1. A method for elucidation of molecular structures from raw liquid chromatography tandem mass spectrometry (LC-MS / MS) data of a sample, comprising: obtaining mass spectrum data, comprising MSI and MS2 data, from the raw LC-MS / MS data, detecting compound peaks by denoising and deconvoluting individual compound fingerprints from the MSI data; predicting a type of charged adduct corresponding to each compound peak represented in the MSI data; determining monoisotopic mass from the predicted charged adduct type, obtaining MSI intensities of an isotope distribution signature for each charged adduct; extracting MS2 spectra for each compound corresponding to a compound peak, based on each detected compound peak, charged adduct and MS 1 intensities; and elucidating a molecular structure by consensus scoring based on at least two of a spectrum score derived by comparing the MS2 spectra to a reference library of reference MS2 spectra and corresponding molecular structures, an isotope score derived by comparing isotope distribution signature to a reference library of calculated isotope distribution signatures of molecular formulae, and a monoisotopic mass comparison score derived by comparing monoisotopic mass of the compound against a reference library of calculated monoisotopic mass of molecular formulae, and extracting the molecular structure from a reference library based on the consensus score.

2. The method of 1, wherein denoising the MSI data comprises generating a graph for each mass-to-charge ratio (m / z) in the MSI data, and categorising each said graph as either compound-like, or noise, and discarding MSI data corresponding to graphs categorised as noise.

3. The method of 2, wherein denoising the mass spectrum data comprises:determining, based on a shape of each graph, whether the respective graph represents a "noise-like" signal, a "fragment-like" signal or a "compound-like" signal; and removing one or both of any said “noise-like” signals and “fragment-like” signals.

4. The method of 3, comprising deconvoluting individual compound fingerprints by characterising each “compound-like” peak based on one or both of an area under curve around a region comprising the peak, and an area under curve outside the region comprising the peak.

5. The method of any one of 1 to 4, wherein charged adducts are predicted based on a summed intensity of all explained adducts relative to a deisotoped MSI spectrum derived from the MSI data, overall agreement of isotope cluster-derived ion charges with predicted adduct ion charges, and a number of explained adducts for a current hypothesis, of one or more hypotheses, relative to a maximum number of explained adducts detected in any of the one or more hypotheses.

6. The method of any one of 1 to 5, obtaining MSI intensities of an isotope distribution signature for each charged adduct comprises extracting the MSI intensities of the isotope distribution signature for each charged adduct, from the MSI data7. The method of any one of 1 to 6, wherein elucidating the molecular structure comprises calculating the spectrum score by applying a weighted cosine similarity to the MS2 spectra for each compound and the reference MS2 spectra, to extract one or more candidate molecular structures from the reference library, wherein weights of the weighted cosine similarity are determined by reference to m / z ratio.

8. The method of 7, wherein elucidating the molecular structure further comprises: the isotope distribution score being a modified cosine similarity score calculated between the isotope distribution signatures from the MSI data, and calculated MSI isotope distributions of molecules in the reference library; determining the molecular structure based on a weighted consensus between each weighted cosine similarity between the MS2 spectra and reference MS2 spectra,each cosine similarity of the isotope distribution signature, and each monoisotopic mass comparison score9. The method of 8, wherein determining the monoisotopic mass comparison score for each candidate molecular structure comprises calculating the respective monoisotopic mass comparison score based on an m / z threshold determined, for each compound, from an m / z of the respective compound, an exact monoisotopic mass of a molecular ion corresponding to the compound peak, and a calculated exact monoisotopic mass of a corresponding molecular ion in the reference library.

10. The method of 9, wherein the weighted MS2 cosine similarity score, modified cosine similarity score and monoisotopic mass comparison score are calculated only for reference spectra within the m / z threshold.

11. A system for elucidation of molecular structures from raw liquid chromatography tandem mass spectrometry (LC-MS / MS) data of a sample, comprising: a receiver for receiving the raw LC-MS / MS data, the raw LC-MS / MS data comprising MSI and MS2 data; a denoising detector for detecting compound peaks by denoising and deconvoluting individual compound fingerprints from the MSI data; an adduct identifier for predicting a type of charged adduct corresponding to each compound peak represented in the MS 1 data; determining monoisotopic mass from the predicted charged adduct type; an isotope characteriser for obtaining MSI intensities of an isotope distribution signature for each charged adduct; a MS2 spectral extractor for extracting MS2 spectra for each compound corresponding to a compound peak, based on each detected compound peak, charged adduct and MS1 intensities; and a molecular analyser for elucidating a molecular structure from the by consensus scoring based on at least two of a spectrum score derived by comparing the MS2 spectra to a reference library of reference MS2 spectra and corresponding molecular structures, an isotope score derived by comparing isotope distribution signature to a reference library of calculated isotope distribution signatures of molecular formulae, and a monoisotopic mass comparison score derived by comparing monoisotopic mass of thecompound against a reference library of calculated monoisotopic mass of molecular formulae, and extracting the molecular structure from a reference library based on the consensus score.

12. The system of 11, wherein the denoising detector denoises the MSI data by generating a graph for each mass-to-charge ratio (m / z) in the MSI data, and categorising each said graph as either compound-like, or noise, and discarding MSI data corresponding to graphs categorised as noise.

13. The system of 12, wherein the denoising detector deconvolutes the mass spectrum data by: determining, based on a shape of each graph, whether the respective graph represents a "noise-like" signal, a "fragment-like" signal or a "compound-like" signal; and removing one or both of any said “noise-like” signals and “fragment-like” signals.

14. The system of 13, wherein the denoising detector characterises each “compound-like” peak based on one or both of an area under curve around a region comprising the peak, and an area under curve outside the region comprising the peak.

15. The system of any one of 11 to 14, wherein the adduct identifier predicts adducts based on a summed intensity of all explained adducts relative to a deisotoped MSI spectrum derived from the MSI data, overall agreement of isotope cluster-derived ion charges with predicted adduct ion charges, and a number of explained adducts for a current hypothesis, of one or more hypotheses, relative to a maximum number of explained adducts detected in any of the one or more hypotheses.

16. The system of any one of 11 to 15, wherein the isotope characteriser obtains MSI intensities of an isotope distribution signature for each charged adduct by extracting the MSI intensities of the isotope distribution signature for each charged adduct, from the MSI data.

17. The system of any one of 11 to 16, wherein the molecular analyser elucidates the molecular structure by calculating the spectrum score by applying a weighted cosine similarity to the MS2 spectra for each compound and the reference MS2 spectra, to extract one or more candidate molecular structures from the reference library, wherein weights of the weighted cosine similarity are determined by reference to m / z ratio.

18. The system of 17, wherein the molecular analyser elucidates the molecular structure further by: calculating the isotope distribution score by calculating a modified cosine similarity score between the isotope distribution signatures from the MSI data, and calculated MSI isotope distributions of molecules in the reference library; determining the molecular structure based on a weighted consensus between each weighted cosine similarity between the MS2 spectra and reference MS2 spectra, each cosine similarity of the isotope distribution signature, and each monoisotopic mass comparison score19. The system of 18, wherein the molecular analyser calculates each said monoisotopic mass comparison score for each candidate molecular structure comprises calculating the respective monoisotopic mass comparison score based on an m / z threshold determined, for each compound, from an m / z of the respective compound, an exact monoisotopic mass of a molecular ion corresponding to the compound peak, and a calculated exact monoisotopic mass of a corresponding molecular ion in the reference library.

20. The system of 19, wherein the molecular analyser is configured to calculate the weighted MS2 cosine similarity score, modified cosine similarity score and monoisotopic mass comparison score only for reference spectra within the m / z threshold.