Structural elucidation using machine learning and alteration of collision energy

By employing machine learning and UMAP for dimensionality reduction of mass spectrometry data, the method addresses the inefficiencies of manual analysis, enabling rapid and precise molecular structure identification in complex mixtures.

WO2025153957A1PCT designated stage expired Publication Date: 2025-07-24DH TECH DEVMENT PTE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
PCT/IB2025/050412
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2025-01-14
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Current methods for determining molecular structures through mass spectrometry are time-consuming and require expert knowledge, especially when dealing with complex mixtures or unknown samples, and lack accurate automatic structural elucidation methods that can rival manual fragment analysis.

Method used

A method using machine learning and dimensionality reduction techniques, specifically Uniform Manifold Approximation & Projection (UMAP), to analyze fragmentation patterns of ions at multiple collision energies, embedding data into a reduced dimension space to identify molecular structures.

Benefits of technology

Enables rapid and effective identification of molecular structures by preserving the underlying structure of high-dimensional data, allowing less skilled personnel to analyze complex mixtures and identify unknown compounds with high precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050412_24072025_PF_FP_ABST
    Figure IB2025050412_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for determining a molecular structure include obtaining mass spectrum data for an unknown parent ion. The mass spectrum data is recorded at each collision energy of a series of collision energies and includes an abundance of one or more substructures of the unknown parent ion. The molecular structure of the unknown parent ion is determined by composing the mass spectrum data into a fragmentation profile and embedding the fragmentation profile into a reduced dimension profile. The reduced dimension profile is mapped into a fragmentation space. One or more proximate molecules in the fragmentation space are identified using the mapping and a candidate for the parent ion is identified using the one or more proximate molecules.
Need to check novelty before this filing date? Find Prior Art

Description

STRUCTURAL ELUCIDATION USING MACHINE LEARNING AND ALTERATION OF COLLISION ENERGYCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application is being filed as a PCT International Patent Application that claims priority to and the benefit of U.S. Provisional Application No. 63 / 621,445, filed on January 16, 2024, the disclosure of which is hereby incorporated by reference in its entirety.BACKGROUND

[0002] Mass spectrometry can be employed in the elucidation of molecular structures by providing valuable information about the mass and fragmentation patterns of ions derived from a given molecule. In the process of structural elucidation, a sample is first subjected to ionization, where it is converted into charged particles. Various ionization techniques, such as electrospray ionization (ESI) or matrix-assisted laser desorption / ionization (MALDI), are utilized depending on the nature of the compound. The resulting ions are then accelerated through an electric or magnetic field, causing them to separate based on their mass-to-charge ratio (m / z). The mass spectrum obtained reflects the distribution of ions, and the molecular ion peak provides the molecular weight of the compound.

[0003] Tandem mass spectrometry (MS / MS) can be employed to unveil further details about a molecule's structure. By selecting a specific ion from the mass spectrum and subjecting it to further fragmentation, the resulting fragment ions and their relative abundances offer insights into the arrangement of atoms within the molecule. This fragmentation pattern aids in identifying functional groups, locating bond positions, and distinguishing isomeric structures. Through the integration of mass spectrometry with other spectroscopic techniques, such as nuclear magnetic resonance (NMR) and infrared spectroscopy, researchers can achieve a comprehensive understanding of a compound's composition and structural characteristics, facilitating its identification in complex mixtures or unknown samples.SUMMARY

[0004] Examples presented herein relate to a method of determining a molecular structure .The method includes obtaining mass spectrum data for an unknown parent ion, the mass spectrum data recorded at each collision energy of a series of collision energies and includingan abundance of one or more substructures of the unknown parent ion; determining the molecular structure of the unknown parent ion by: composing the mass spectrum data into a fragmentation profde; embedding the fragmentation profde into a reduced dimension profde; mapping the reduced dimension profde into a fragmentation space; identifying one or more proximate molecules in the fragmentation space using the mapping; and identifying a candidate for the parent ion using the one or more proximate molecules.

[0005] In other aspects presented herein, embedding the fragmentation profde into a reduced dimension profde comprises applying a Uniform Manifold Approximation & Projection (UMAP) algorithm to the fragmentation profde. In further aspects presented herein, the reduced dimension profde comprises a heatmap. In other further aspects presented herein, the reduced dimension profde has two dimensions.

[0006] In other aspects presented herein, determining the molecular structure of the parent ion is performed by a machine learning model. In further aspects presented herein, the machine learning model is trained by: obtaining high dimensional mass spectra data for a plurality of known compounds; embedding, for each of the plurality of known compounds, the high dimensional mass spectra data into a reduced dimension profde; providing each reduced dimension profde to the machine learning model for combining into the fragmentation space. In yet further aspects presented herein, the high dimensional mass spectra data comprises one or more fragmentation curves for substructures of the known compound. In still further aspects presented herein, embedding the high dimensional mass spectra data into a reduced dimension profde comprises applying a Uniform Manifold Approximation & Projection (UMAP) algorithm to the heatmap. In yet further aspects, the known compounds are small molecules.

[0007] In other aspects presented herein, the series of collision energies includes at least two collision energies, at least ten collision energies, or at least 16 collision energies. In yet other aspects presented herein, the series of collision energies spans from 5 - 100 eV.

[0008] In other aspects presented herein, each of substructure of the one or more substructures is identified as being present at one or more collision energies of a series of collision energies. In further aspects presented herein, each of substructure of the one or more substructures is identified as being present at two or more collision energies of a series of collision energies. In yet further aspects presented herein, each of substructure of the one or more substructures is identified as being present at five or more collision energies of a series of collision energies. In still further aspects presented herein, the collision energy, of the oneor more collision energies, which has a highest intensity for the substructure is identified and added to the fragmentation profile.

[0009] In other aspects presented herein, the fragmentation profile comprises a dataframe spanning the mass range 50 - 550 Da. In still other aspects presented herein, the reduced dimension profile and the fragmentation space have a same number of dimensions. In further aspects presented herein, the same number of dimensions is two dimensions.

[0010] In other aspects presented herein, embedding the fragmentation profile into a reduced dimension profile comprises applying one of a PCA and a tSNE algorithm to the fragmentation profile. In yet other aspects presented herein, at least of the one or more proximate molecules is a nearest neighbor and identifying a candidate for the parent ion using one or more proximate molecules comprises using the nearest neighbor.

[0011] Other examples presented herein relate to a system for determining a molecular structure of a parent ion. The system including a processor and a non-transitory memory in communication with the processor and storing instructions that, when executed, cause the processor to: obtain mass spectrum data for an unknown parent ion, the mass spectrum data recorded at each collision energy of a series of collision energies and including an abundance of one or more substructures of the unknown parent ion; determine the molecular structure of the unknown parent ion by: composing the mass spectrum data into a fragmentation profile; embedding the fragmentation profile into a reduced dimension profile; mapping the educed dimension profile into a fragmentation space and identifying a nearest neighbor in fragmentation space; and identifying a candidate for the parent ion using the nearest neighbor.

[0012] Other examples presented herein relate to a method of training a machine learning model for determining a molecular structure. The method including obtaining high dimensional mass spectra data for a plurality of known compounds, wherein, for each of the plurality of known compounds, the high dimensional mass spectra data is recorded at each collision energy of a series of collision energies and includes an abundance of one or more substructures of the known compound; composing, for each of the plurality of known compounds, a fragmentation profile using the high dimensional mass spectra data; embedding, for each of the plurality of known compounds, the high dimensional mass spectra data into a reduced dimension profile; and providing each reduced dimensionality fragmentation profile to the machine learning model for combining into a common fragmentation space.

[0013] In other aspects presented herein, once trained, the machine learning model infers the molecular structure of an unknown compound by fitting a fragmentation profile of theunknown compound into the common fragmentation space. In yet other aspects presented herein, the common fragmentation space has two dimensions. In still other aspects presented herein, the fragmentation profde is configured as a heatmap. In other aspects, the known and unknown compounds are small molecules and do not include peptides or nucleic acid chains.

[0014] A variety of additional inventive aspects will be set forth in the description that follows. The inventive aspects can relate to individual features and to combinations of features. It is to be understood that both the forgoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the broad inventive concepts upon which the embodiments disclosed herein are based.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings, which are incorporated in and constitute a part of the description, illustrate several aspects of the present disclosure. A brief description of the drawings is as follows:

[0016] FIG. 1 is a block diagram of an example system for determining a structure or identify of an unknown compound.

[0017] FIG. 2 is block diagram of an example data processing system for determining a molecular structure or identify.

[0018] FIG. 3 is a flowchart of a method of determining a molecular structure.

[0019] FIG. 4 is a flowchart of a method for determining the molecular structure of an unknown parent ion.

[0020] FIG. 5 is an example process flow of generating stacked fragmentation curves from mass spectrum data.

[0021] FIG. 6 is a graphic illustration of some aspects of the method of FIG. 4.

[0022] FIG. 7 is an example of a fragmentation profde depicted as a heatmap.

[0023] FIG. 8 is an example of a reduced dimension profde for an unknown being mapped into the fragmentation space.

[0024] FIG. 9 is a flowchart of a method for training a machine learning model for determining a molecular structure.

[0025] FIG. 10 illustrates an example block diagram of a virtual or physical computing system.DETAILED DESCRIPTION

[0026] Ions fragment in the mass spectrometer in a way that is determined in large part by their molecular structure and the collision energy (CE) setting of the instrument. During the fragmentation process, molecules (e.g., small molecules) break down into the most stable units, many of which are substructures shared between similar as well as disparate molecular classes (e.g., benzene rings). The combination of substructures derived from a molecule are a fingerprint of the structure (or a set of very closely related structures) and can be analyzed either manually or automatically to elucidate the structure of the parent ion.

[0027] Performed manually by an expert, determining structure from fragmentation patterns is time consuming and incompatible with high-throughput work. However, no automatic structural elucidation method has been developed that rivals the precision achieved by manual fragment analysis. Compound identification from tandem mass spectra (MS / MS) and mass spectra generally is extremely challenging when there is not an exact match in an existing spectral library. Users spend considerable time on compound identification and expertise is generally required. Even moderate improvements over the current state of the art will offer considerable benefits, by allowing less skilled personnel to obtain useful results in a majority of cases.

[0028] The absence of an accurate automatic structural elucidation arises from the challenge presented by defined substructures being present in disparate molecular classes. Additional dimensionality in spectral data is needed to elucidate true fragment trees from mass spectra. Disclosed herein are novel systems and methods for identification of molecules based on signature fragmentation which provide for more rapid and effective identification of a structure, and in some cases an identity, of a parent ion based on a fragmentation pattern generated by the parent ion.

[0029] Machine learning (ML) is well suited to solve the problem of determining molecular structure from mass spectra because it excels at sensing patterns in highdimensional, complex data. Instead of hard-coded algorithms, which are time consuming to code and relatively inflexible, a ML approach establishes a model with many parameters (e.g., 103- 109) that are tuned by feeding training data to the model. In the case of mass spectra data, training data generally consists of experimentally measured mass spectra paired to the identity of the molecule that generated them, serving as examples from which the model can learn to predict unknown molecular structure from spectra.

[0030] Judicious design of the training data is crucial to the development of the ML model .As disclosed herein, the spectra and molecules are represented in a new way to present as much information as possible to the model. Separate spectra for each molecule are acquired at a multitude of CEs to provide insight into the fragmentation pathways of a molecule. In embodiments, the data is acquired in a continuous ramp mode. Instead of being averaged, the spectra from different CEs are then be stacked into a 2D image representing information on the collision induced fragmentation of a single molecule.

[0031] Chemical space and molecular structures, like those used to generate the training data, are most commonly represented in informatics settings as strings, for example as Simplified Molecular Input Line Entry System (SMILES) or International Chemical Identifier (InChi). However these one-dimensional strings do not directly represent the three-dimensional structure of molecules, and require the ML model to learn their idiosyncratic grammar. As is disclosed herein, the molecules associated with each set of spectra are represented in multiple dimensions to allow the model to incorporate a higher degree of molecular structural information.

[0032] In order to generate a manageable training data set from the large amount of complex and high-dimensional data generated, dimensionality reduction is employed. Dimensionality reduction is a powerful tool for highlighting patterns in data that have too many dimensions to be readily visualized. For example, a blood sample may be analyzed for 100 different biomarkers, then reduced to one dimension: probability of a particular disease. One algorithmic implementation of dimensionality reduction is Uniform Manifold Approximation & Projection (UMAP). UMAP takes data with, for example, 102- 105dimensions and embeds the data into a lower number of dimensions, e.g., 2 - 102dimensions. Importantly, in the lowerdimensional space, UMAP preserves both local and global structure. In other words, points that are close together in the embedding share patterns that UMAP has detected in the higherdimensional space.

[0033] While advantages of the present disclosure and examples presented herein are described in terms of UMAP, other dimensionality reduction algorithms may be applied under the same principles. For example, one or more oft-Distributed Stochastic Neighbor Embedding (t-SNE), Principal Component Analysis (PCA), Locally Linear Embedding (LLE), Isometric Mapping (ISOMAP), autoencoders, Potential of Heat-diffusion for Affinity-based Trajectory Embedding (PHATE), diffusion maps, or Laplacian Eigenmaps may be appropriate fordimension reduction in some implementations. In embodiments, the CE information may be reduced to vectors (e.g., increasing vs decreasing) and used as input to ML.

[0034] UMAP and other dimensionality reduction algorithms are primarily used to reduce the dimensionality of data while preserving its underlying structure. It projects highdimensional data points into a lower-dimensional space, typically two-dimensional or three- dimensional, where the relationships between data points are preserved as much as possible. UMAP is also designed to capture the underlying manifold structure of data. Manifolds are lower-dimensional representations of complex, high-dimensional data that capture the intrinsic relationships between data points. UMAP aims to find a representation that respects the topology and distances between points on the manifold. UMAP can uncover complex, nonlinear patterns in the data, making it suitable for a wider range of datasets.

[0035] As disclosed herein, high-dimensional data is generated from mass spectra of known compounds (e.g., small molecules) collected at different CEs. In embodiments, mass spectra may be collected at 2, 3, 4, 6, 10, 15, 16, 18, 20, etc. different CEs. In embodiments, the different CEs may cover a range, for example, of 5 - 100 eV. For each compound, fragments are identified by finding m / z values that appear at multiple CEs. In embodiments, fragments may be identified by appears at least two, at least three, at least five, at least ten, etc. difference CEs. In a particular example, fragments may be identified by appearing at least five different CEs.

[0036] For some or each of these fragments, a CE at which the fragment has the highest intensity is identified and added to a dataframe. In embodiments, the dataframe may span a mass range of 50 - 550 Da at the appropriate m / z point. In an example using nominal mass, a 501 dimension dataset results, however those of skill in the will understand this could be expanded to take advantage of high-resolution mass data. An example table for one compound is shown in Table 1 ; an entry of 0 indicates that the compound does not form a fragment at that mass:TABLE 1

[0037] According to the present disclosure, a UMAP ML model is trained on the 501- dimensional data and learns to embed this data into a lower dimensional space, e.g., into two dimensions. The model is subsequently applied to inference the structure of an unknowncompound, which is collected in the same high dimensional format as the training data. The unknown is embedded into the same two-dimensional space by the model. Inferences about the unknown are drawn based on distance and relationship to known compounds in the two- dimensional space. For example, the unknown’s nearest neighbors in the two-dimensional embedding are used to identify likely substructures present in the unknown. In embodiments, likely substructures may be identified using chemoinformatic tools such as RDKIT.

[0038] The systems and methods disclosed herein present a number of advantages over existing processes for molecular identification using multiple collision energies. For example, a user is able to view a visualization of a degree of “similarity” of a series of three-dimensional structures of the molecules, with similarity determined based through the closeness of the CE profiles of MS / MS fragments of the molecules. This multi- dimensional data representation reveals the relationships in the chemical space as a whole, capturing the relationships between the compounds within the chemical space. This representation offers more information and is superior to the classical “unknown identification reporting” that generally identifies many references compounds in response to an unknown. Further, the proposed systems and method do not rely on library searching and therefore can be used to identify molecules that have not been catalogued. By connecting the collision energy to substructure masses, then allowing a model to find high-dimensional patterns, large amounts of quality information about molecular structure is incorporated into the analysis. Unknown compound structural elucidation aides users in identifying new metabolites in complex mixtures, or screen for unknown substances in a high throughput fashion.

[0039] Particular to embodiments incorporating UMAP, UMAP is often faster when compared to other dimensionality reduction algorithms, and embedding a new spectrum withing an existing model can be performed in a matter of seconds (e.g., ~2 seconds in some cases) on a standard processing device.

[0040] FIG. 1 is a block diagram of an example system 100 for determining a structure or identify of an unknown compound. In embodiments, systems for determining compound structure and identity may be integrated with a mass spectrometry system, as shown in example system 100, or, in other embodiments, may be independent of the mass spectrometry system. Example system 100 includes an ion source 110, a first mass separator 120, a fragmentation device 130, a second mass separator or a mass analyzer 140, and a computing system 150.

[0041] In embodiments, system 100 further includes a sample introduction device 160. Sample introduction device 160 introduces one or more compounds of interest from a sampleto ion source 110 overtime. Sample introduction device 160 performs techniques that include, but are not limited to, direct injection, liquid chromatography, gas chromatography, capillary electrophoresis, or ion mobility.

[0042] Mass fdter 120 and fragmentation device 130 are shown as different stages of a quadrupole and mass analyzer 140 is shown as atime-of-flight (TOF) device. Those of ordinary skill in the art will appreciate that either of mass filter 120 and mass analyzer 140 may include other types of mass separator and analysis devices including, but not limited to, ion traps, orbitraps, ion mobility devices, time-of-flight (TOF) devices, or Fourier transform ion cyclotron resonance (FT-ICR) devices. In embodiments, mass filter 120 and mass analyzer 140 are respective examples of a first and a second mass separator, arranged in a series. A system may be configured according to the present disclosure with a quadrupole for the first mass separator, or mass filter 120, and a TOF device for the second mass separator, or mass analyzer 140. Each mass separator is configured to receive a set of ions, perform a detection of the set of ions, and generate a set of detection signals corresponding to detection of the set of ions.

[0043] Ion source device 110 transforms a sample or compounds of interest from a sample into an ion beam. Ion source device 110 can perform ionization techniques that include, but are not limited to, matrix assisted laser desorption / ionization (MALDI) or electrospray ionization (ESI).

[0044] Mass filter 120 receives the ion beam. In embodiments, mass filter 120 is configured by a user for a particular precursor ion transmission window based on the experimental goals for the sample being run. The precursor ion transmission window, as discussed herein, refers to the range of precursor or parent ions that are allowed to pass through a specific selection step and into the subsequent stages of mass analysis or fragmentation. In many tandem mass spectrometry (MS / MS) experiments, the precursor ions are first selected based on their m / z (mass-to-charge ratio) in order to isolate a specific ion of interest for further analysis or fragmentation. The precursor ion selection process employs a mass filter or a specific set of voltages that allow only ions within a certain m / z range (the precursor ion transmission window) to pass through to the next stage.

[0045] Fragmentation device 130 of tandem mass spectrometer 102 fragments or transmits the precursor ions transmitted by mass filter 120. In examples related to data independent acquisition, including scanning SWATH specifically, one or more resulting product ions are produced for each overlapping window of the series. Fragmentation device 130 fragments the precursor ions when a collision energy high enough to fragment ions is used. Fragmentationdevice 130 transmits the precursor ions when a collision energy low enough not to fragment ions is used. As a result, the resulting product ions can include precursor ions.

[0046] Collision energy influences the fragmentation of ions during collision-induced dissociation (CID) or collision-induced fragmentation (CIF). This process is commonly used in MS / MS to provide structural information about a molecule by inducing the dissociation of parent or precursor ions into fragment ions. The collision energy is the kinetic energy imparted to the precursor ions during collisions with a collision gas, such as helium or nitrogen, in the collision cell of the mass spectrometer.

[0047] Collision energy may be controlled by the instrument operator and can be adjusted to optimize the fragmentation of ions for a specific analysis. The collision energy is often expressed in electronvolts (eV) or joules (J) and is specific to the instrument and the type of collision cell used. The setting of the collision energy impacts the resulting fragmentation pattern and may influence how meaningful any acquired data is. For a single experimental rule, collision energy being too low may result in insufficient fragmentation, while too high a collision energy can lead to excessive fragmentation, causing the loss of important structural information.

[0048] By systematically increasing or decreasing the collision energy, researchers can observe how the fragmentation pattern changes. This approach, which may be referred to as collision energy ramping, allows for the identification of the optimal collision energy that yields the most informative and comprehensive structural information for a particular compound. Collision energy ramping may also be applied to develop a fragmentation profile of a parent or precursor ion. By identifying a precursor ion based on characteristic patterns across multiple collision energies, rather than a single CE, greater structural data is acquired.

[0049] Mass analyzer 140 of tandem mass spectrometer 102 detects intensities or counts for each of the one or more resulting product ions for each overlapping window of the series that form mass spectrum data for each overlapping window of the series.

[0050] Computing system 150 can be, but is not limited to, a computer, a microprocessor, the computing system of FIG. 7, or any device capable of sending and receiving control signals and data from a tandem mass spectrometer and processing data. Computing system 150 is in communication with ion source device 110, mass filter 120, fragmentation device 130, and mass analyzer 140. Computing system 150 is shown as a separate device but can be a processor or controller of tandem mass spectrometer 102 or another device. Computing system 150 may store in a memory device (not shown) mass spectrum data for each precursor ion windowanalysis is performed for, including for each overlapping window of the series in examples performing scanning SWATH. In embodiments, computing system 150 instead performs an encoding and storing step, and encodes and stores each unique product ion detected by mass analyzer 140 in real-time during data acquisition. Prior to storing mass spectrum data, computing system 150 performs one or more processing steps on the raw mass spectrum data received to prepare the data for viewing, analysis, and storage. Raw mass spectrum data includes the counts or intensities of product ions at different m / z ratios over time.

[0051] FIG. 2 is block diagram of an example data processing system 200 for determining a molecular structure or identify. Data processing system 200 is implemented, in embodiments, by computing system 150. Data processing system 200 may be among multiple subsystems or software executed by computing system 150 in operating mass spectrometer 102 and handling of data output by the mass spectrometer. The present disclosure is directed to determining molecular structure or identity, but those of skill in the art will readily understand that other subsystems and / or modules may exist within and be executed by computing system 150. Computing system 150 is presented in examples herein as a single device, but in embodiments may be one or more processing devices networked or otherwise in communication. Functions may be divided among individual devices or shared across the collective processing capability of the one or more processing devices. In embodiments, data processing system 200 forms part of or serves as a controller. The controller may be configured to transmit operation commands to the mass spectrometer and / or to receive unprocessed mass data from second mass separator including the set of detection signals and perform one more data processing actions on the unprocessed mass data.

[0052] In the example of FIG. 2, data processing system 200 includes preprocessor 204, dimension reduction 206, fragmentation model 208, fragmentation profile 210, and file formatter 212.

[0053] Raw mass spectrometer data is typically large and complex, and it undergoes extensive data processing and analysis to extract meaningful information. Data processing may be carried out in a series or group of actions. Actions may be performed collectively by the data processing system 200 or individual steps or portions of the processing may be executed by individual components or modules of the data processing system 200. In some embodiments, some components or features shown as integrated with data processing system 200 may instead or in addition be executed elsewhere on an external component.

[0054] Preprocessor 204 performs one or more processing operations on mass spectrometry data 202 to prepare the data for indexing, storage, and subsequent analysis. Before indexing, the mass spectrometry data often undergoes preprocessing, which includes, by example, data conversion, noise reduction, peak picking, and deconvolution. This step helps simplify the data and enhances the quality of the information to be indexed. Peaks, representing ions and their corresponding intensities at specific m / z values, are identified and quantified. The detected peaks are typically used as the basis for indexing.

[0055] Dimension reduction 206 executes one or more dimension reduction algorithms on the mass spectra data 202. In embodiments, dimension reduction 206 includes applying a UMAP algorithm to the high dimensional mass spectra data. As discussed above, UMAP may provide some advantages in some embodiments, but other dimension reduction algorithms will also be applicable and appropriate to embodiments of the present disclosure. Other dimensionality reduction algorithms may be applied under the same principles. For example, one or more of t-Distributed Stochastic Neighbor Embedding (t-SNE), Principal Component Analysis (PCA), Locally Linear Embedding (LLE), Isometric Mapping (ISOMAP), autoencoders, Potential of Heat-diffusion for Affinity-based Trajectory Embedding (PHATE), diffusion maps, or Laplacian Eigenmaps may be appropriate for dimension reduction in some implementations.

[0056] Fragmentation model 208 is a machine learning model trained to analyze a fragmentation pattern for an unknown compound and map the compound into a fragmentation space. In embodiments, fragmentation model 208 may incorporate aspects of dimension reduction 206.

[0057] Fragmentation model 208 may be trained by embedding, for each of a plurality of known compounds, high dimensional mass spectra data into a reduced dimension profile. The high dimensional mass spectra data, for each of the plurality of known compounds, may include one or more fragmentation curves for substructures of the known compound. The mass spectra data may be collected at a series of collision energies for each of the plurality of known compounds. In embodiments, the series of collision energies includes at least two collision energies, at least ten collision energies, at least 16 collision energies, etc. In some cases, the series of collision energies spans from 5 - 100 eV, though other suitable ranges for the series of collision energies are envisioned and may be applied by those of skill in the art as dictated by experimental guidelines in various use cases.

[0058] Fragmentation space 210 is the stored fragmentation space generated by fragmentation model 208 and then used to draw inferences about unknown compounds analyzed. Fragmentation space 210 may be generated once and used for multiple inferences. In embodiments, fragmentation space 210 may be generated based on a curated profde of known compounds for a particular experimental run or series of experimental runs.

[0059] File formatter 212 converts the final processed mass spectrum data into a standardized file format for storage and transmittal. In embodiments, the standardized file format is a wiff file.

[0060] FIG. 3 is a flowchart of a method 300 of determining a molecular structure. In embodiments, the method 300 is performed by a data processing system, such as data processing system 200 of FIG. 2.

[0061] At operation 302, mass spectrum data for an unknown parent ion is obtained. The mass spectrum data includes an abundance of one or more substructures of the unknown parent ion. Mass spectra data generally provides information about the mass-to-charge ratios of substructures of the parent ion produced, offering a fingerprint of a compound's molecular composition and structure through the unique pattern of peaks and fragment ions. In embodiments, obtaining the mass spectrum data involves receiving a processed data file from a mass spectrometer or a storage system. In some cases, obtaining the mass spectrum data includes collecting the data by the mass spectrometer.

[0062] The mass spectrum data is recorded at each collision energy of a series of collision energies. In embodiments, the series of collision energies includes at least two collision energies, at least ten collision energies, at least 16 collision energies, etc. In some cases, the series of collision energies spans from 5 - 100 eV, though other suitable ranges for the series of collision energies are envisioned and may be applied by those of skill in the art as dictated by experimental guidelines in various use cases.

[0063] At operation 304, the molecular structure of the unknown parent ion is determined. In embodiments, determining the molecular structure of the unknown parent ion may encompass identifying the parent ion.

[0064] FIG. 4 is a flowchart of a method 400 for determining the molecular structure of an unknown parent ion. The method 400 may be a subprocess of step 304 of method 300 and represents the steps taken by the system to determine the molecular structure of a parent ion. In embodiments, the method 400 is performed by a machine learning model. In some cases,the machine learning model may be a subcomponent of a data processing system executing method 300, such as fragmentation model 208 of data processing system 200 of FIG. 2.

[0065] The machine learning model may be trained by embedding, for each of a plurality of known compounds, high dimensional mass spectra data into a reduced dimension profile. The high dimensional mass spectra data, for each of the plurality of known compounds, may include one or more fragmentation curves for substructures of the known compound. The mass spectra data may be collected at a series of collision energies for each of the plurality of known compounds. In embodiments, the series of collision energies includes at least two collision energies, at least ten collision energies, at least 16 collision energies, etc. In some cases, the series of collision energies spans from 5 - 100 eV, though other suitable ranges for the series of collision energies are envisioned and may be applied by those of skill in the art as dictated by experimental guidelines in various use cases.

[0066] In embodiments, embedding the high dimensional mass spectra data into a reduced dimension profde includes applying a UMAP algorithm to the high dimensional mass spectra data. As discussed above, UMAP may provide some advantages in some embodiments, but other dimension reduction algorithms will also be applicable and appropriate to embodiments of the present disclosure. Each, or some portion, of the reduced dimension profiles are then provided to the machine learning model. The model then combines the reduced dimension profiles into a fragmentation space.

[0067] At operation 402, the mass spectrum data is composed into a fragmentation profile . The high dimensional mass spectra data comprises one or more fragmentation curves for substructures of the unknown compound. Fragmentation curves represent the relationship between the intensity of ion fragments and the collision energy applied during collision- induced dissociation (CID) or other fragmentation processes. These curves illustrate how the abundance of specific fragment ions changes as a function of the collision energy. According to the present disclosure, the fragmentation curves can be stacked into a fragmentation profile, instead of being averaged together or considered separately. In this way, multiple fragmentation curves from multiple CEs can be combined into a wholistic profile, without the loss of individual features of the various CEs. FIG. 5 is an example process flow 500 of generating stacked fragmentation curves from mass spectrum data.

[0068] The fragmentation profile may be generated as a dataframe, such as the example presented in Table 1 above. For example, the fragmentation profile in some embodiments includes a dataframe spanning the mass range 50 - 550 Da.

[0069] In embodiments, each of substructure of the one or more substructures is identified as being present at one or more collision energies of a series of collision energies, at two or more collision energies of a series of collision energies, at five or more collision energies of a series of collision energies, etc. For example, in some cases, each substructure produced may be recorded into the fragmentation profile, orthose appearing at only one CE (or only two CEs, or less than 5 CEs,) may be discarded to simplify the data and provide clearer results.

[0070] In embodiments, the collision energy, of the one or more collision energies, which has a highest intensity for the substructure is identified and added to the fragmentation profile. For example, once a substructure is identified as being present at a sufficient number of CEs, it is recorded into the fragmentation profile at the CE for which the substructure produces the highest intensity. This collision energy into the dataframe cell corresponding to the compound and the fragment mass. An example is presented in Table 2 below.TABLE 2

[0071] In the example of Table 2, a “0” indicates a compound produces no fragment at a particular mass. For mass at which a compound does produced a fragment, the CE of the most intense appearance is shown. For example, Compound 1 produces a fragment with a mass of 120 Da which is most intense at 21 eV and another with a mass of 135 Da which is most intense at 8 eV. In some examples, only fragments between 50-550 Da are considered, but other mass ranges are contemplated and may be applicable given experimental parameters of some applications of the present disclosure. This example data frame, or fragmentation profile, is considered to 501 dimensions, because each entry is described by 501 variables.

[0072] FIG. 6 is a process flow of aspects of the method 400 of FIG. 4. A plurality of fragmentation curves 452 are combined into a fragmentation profile 454. In embodiments, the fragmentation profile 454 is depicted as a heatmap. In some cases, the reduced dimension profile has two dimensions.

[0073] FIG. 7 is an example of a fragmentation profile depicted as a heatmap. As shown in this example, the fragmentation curve data for a single molecule can be represented as a heatmap. In this example, the x position represents mass where each vertical streak is a fragment and the y position represents collision energy. The colour of each pixel may represent the normalized intensity of the measurement for that mass / CE pair. Vertical streaks corresponding to each fragment may not be continuous because data is not collected at every possible collision energy.

[0074] At operation 404, the fragmentation profile is embedded into a reduced dimension profile. In embodiments, embedding the fragmentation profile into a reduced dimension profile includes applying a UMAP algorithm to the fragmentation profile. In some cases, embedding the fragmentation profile into a reduced dimension profile is accomplished by applying one of a PCA and a tSNE algorithm to the fragmentation profile.

[0075] At operation 406, the reduced dimension profile is mapped into a fragmentation space. As can be seen in FIG. 6, the reduced dimension profile is received by ML model 456 and placed within fragmentation space 458. In some preferred embodiments, the reduced dimension profile and the fragmentation space have a same number of dimensions, which may be two dimensions.

[0076] At operation 408, one or more proximate molecules are identified in the fragmentation space using the mapping. FIG. 8 is an example of a reduced dimension profile for an unknown being mapped into the fragmentation space. In this example, unknown 802 is mapped into fragmentation space 804 by a UMAP model. An area of proximity 806 around the unknown 802 may be used to identify one or more neighbors 808. In embodiments, at least of the one or more proximate molecules is a nearest neighbor and identifying a candidate for the parent ion using one or more proximate molecules comprises using the nearest neighbor.

[0077] At operation 410, a candidate for the parent ion is identified using the one or more proximate molecules. Each of the one or more neighbors 808 has a known associated structure, which may be applied by a user or the system to infer a structure of the unknown 802. A trained expert or machine learning model can use the structures of the proximate known molecules to generate candidate structures for the unknown compound.

[0078] FIG. 9 is a flowchart of a method 500 for training a machine learning model for determining a molecular structure. At operation 502, high dimensional mass spectra data for a plurality of known compounds is obtained. The high dimensional mass spectra data is recorded, for each of the known compounds, at each collision energy of a series of collision energies.The mass spectra data generally includes an abundance of one or more substructures of the known compound.

[0079] At operation 504, a fragmentation profde is composed, using the high dimensional mass spectra data, for each of the known compounds. In some cases, the fragmentation profde is configured as a heatmap or a dataframe. At operation 506, the high dimensional mass spectra data is embedded, from each fragmentation profile, into a reduced dimension profile.

[0080] At operation 508, each reduced dimensionality fragmentation profile is provided to the machine learning model to be trained. As the model evaluates the received training data, patterns are identified and each known compound is combined into a common fragmentation space. In some cases, the common fragmentation space has two dimensions. This common fragmentation space can subsequently be used to make inferences about the structure and identity of unknown compounds evaluated using the model. In embodiments, the machine learning model infers the molecular structure of an unknown compound by fitting a fragmentation profile of the unknown compound into the common fragmentation space.

[0081] FIG. 10 illustrates an example block diagram of a virtual or physical computing system 150. One or more aspects of the computing system 150 can be used to implement the systems and methods for constructing data structures for high-dimensionality extraction. In particular, the computing system 150 may be used to implement the data processing system 200 and underlying or integrated components, such as dimension reduction 206 and fragmentation model 208.

[0082] In the embodiment shown, the computing system 150 includes one or more processors 152, a system memory 158, and a system bus 172 that couples the system memory 158 to the one or more processors 152. The system memory 158 includes RAM (Random Access Memory) 160 and ROM (Read-Only Memory) 162. A basic input / output system that contains the basic routines that help to transfer information between elements within the computing system 150, such as during startup, is stored in the ROM 162. The computing system 150 further includes a mass storage device 164. The mass storage device 164 is able to store software instructions and data. Mass storage device 164 may correspond to storage device 212 of FIG. 2, and store index 214 and matrices 216 as described in greater detail above. The one or more processors 152 can be one or more central processing units or other processors.

[0083] The mass storage device 164 is connected to the one or more processors 152 through a mass storage controller (not shown) connected to the system bus 172. The mass storage device 164 and its associated computer-readable data storage media providenonvolatile, non-transitory storage for the computing system 150. Although the description of computer-readable data storage media contained herein refers to a mass storage device, such as a hard disk or solid state disk, it should be appreciated by those skilled in the art that computer-readable data storage media can be any available non-transitory, physical device or article of manufacture from which the central display station can read data and / or instructions.

[0084] Computer-readable data storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable software instructions, data structures, program modules or other data. Example types of computer-readable data storage media include, but are not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid state memory technology, CD-ROMs, DVD (Digital Versatile Discs), other optical storage media, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the computing system 150.

[0085] According to various embodiments of the invention, the computing system 150 may operate in a networked environment using logical connections to remote network devices through the network 148. The network 148 is a computer network, such as an enterprise intranet and / or the Internet. The network 148 can include a LAN, a Wide Area Network (WAN), the Internet, wireless transmission mediums, wired transmission mediums, other networks, and combinations thereof. The computing system 150 may connect to the network 148 through a network interface unit 154 connected to the system bus 172. It should be appreciated that the network interface unit 154 may also be utilized to connect to other types of networks and remote computing systems. The computing system 150 also includes an input / output controller 156 for receiving and processing input from a number of other devices, including a touch user interface display screen, or another type of input device. Similarly, the input / output controller 156 may provide output to a touch user interface display screen or other type of output device.

[0086] As mentioned briefly above, the mass storage device 164 and the RAM 160 of the computing system 150 can store software instructions and data. The software instructions include an operating system 168 suitable for controlling the operation of the computing system 150. The mass storage device 164 and / or the RAM 160 also store software instructions, that when executed by the one or more processors 152, cause one or more of the systems, devices, or components described herein to provide functionality described herein. For example, the mass storage device 164 and / or the RAM 160 can store software instructions that, whenexecuted by the one or more processors 152, cause the computing system 150 to receive and execute managing network access control and build system processes.

[0087] The software instructions further include one or more software applications 166. Software applications may include dedicated systems and algorithms for performing specific tasks or actions or providing specific interfaces. One or more of data processing system 200 and / or one or more component of data processing system 200 may be encompassed by software applications 166.

[0088] Illustrative examples of the systems and methods described herein are provided below. An embodiment of the system or method described herein may include any one or more, and any combination of, the clauses described below.

[0089] Clause 1. A method of determining a molecular structure, including obtaining mass spectrum data for an unknown parent ion, the mass spectrum data recorded at each collision energy of a series of collision energies and including an abundance of one or more substructures of the unknown parent ion; determining the molecular structure of the unknown parent ion by: composing the mass spectrum data into a fragmentation profile; embedding the fragmentation profile into a reduced dimension profile; mapping the reduced dimension profile into a fragmentation space; identifying one or more proximate molecules in the fragmentation space using the mapping; and identifying a candidate for the parent ion using the one or more proximate molecules.

[0090] Clause 2. The method of clause 1, wherein embedding the fragmentation profile into a reduced dimension profile includes applying a Uniform Manifold Approximation & Projection (UMAP) algorithm to the fragmentation profile.

[0091] Clause 3. The method of clause 2, wherein the reduced dimension profile includes a heatmap.

[0092] Clause 4. The method of clause 2, wherein the reduced dimension profile has two dimensions.

[0093] Clause 5. The method of any one of clauses 1-4, wherein determining the molecular structure of the parent ion is performed by a machine learning model.

[0094] Clause 6. The method of clause 5, wherein the machine learning model is trained by: obtaining high dimensional mass spectra data for a plurality of known compounds; embedding, for each of the plurality of known compounds, the high dimensional mass spectra data into a reduced dimension profile; providing each reduced dimension profile to the machine learning model for combining into the fragmentation space.

[0095] Clause 7. The method of clause 6, wherein the high dimensional mass spectra data includes one or more fragmentation curves for substructures of the known compound.

[0096] Clause 8. The method of clause 6, wherein embedding the high dimensional mass spectra data into a reduced dimension profile includes applying a Uniform Manifold Approximation & Projection (UMAP) algorithm to the heatmap.

[0097] Clause 9. The method of any one of clauses 1-8, wherein the series of collision energies includes at least two collision energies.

[0098] Clause 10. The method of clause 9, wherein the series of collision energies includes at least ten collision energies.

[0099] Clause 11. The method of clause 10, wherein the series of collision energies includes at least 16 collision energies.

[0100] Clause 12. The method of any one of clauses 1-11, wherein the series of collision energies spans from 5 - 100 eV.

[0101] Clause 13. The method of any one of clauses 1-12, wherein each of substructure of the one or more substructures is identified as being present at one or more collision energies of a series of collision energies.

[0102] Clause 14. The method of clause 13, wherein each of substructure of the one or more substructures is identified as being present at two or more collision energies of a series of collision energies.

[0103] Clause 15. The method of clause 14, wherein each of substructure of the one or more substructures is identified as being present at five or more collision energies of a series of collision energies.

[0104] Clause 16. The method of clause 13, wherein the collision energy, of the one or more collision energies, which has a highest intensity for the substructure is identified and added to the fragmentation profile.

[0105] Clause 17. The method of any one of clauses 1-16, wherein the fragmentation profile includes a dataframe spanning the mass range 50 - 550 Da.

[0106] Clause 18. The method of any one of clauses 1-17, wherein the reduced dimension profile and the fragmentation space have a same number of dimensions.

[0107] Clause 19. The method of clause 18, wherein the same number of dimensions is two dimensions.

[0108] Clause 20. The method of any one of clauses 1-19, wherein embedding the fragmentation profde into a reduced dimension profde includes applying one of a PC A and a tSNE algorithm to the fragmentation profile.

[0109] Clause 21. The method of any one of clauses 1-20, wherein at least one of the one or more proximate molecules is a nearest neighbor and identifying a candidate for the parent ion using one or more proximate molecules comprises using the nearest neighbor.

[0110] Clause 22. A system for determining a molecular structure of a parent ion, including a processor; a non-transitory memory in communication with the processor and storing instructions that, when executed, cause the processor to: obtain mass spectrum data for an unknown parent ion, the mass spectrum data recorded at each collision energy of a series of collision energies and including an abundance of one or more substructures of the unknown parent ion; determine the molecular structure of the unknown parent ion by: composing the mass spectrum data into a fragmentation profile; embedding the fragmentation profile into a reduced dimension profile; mapping the educed dimension profile into a fragmentation space and identifying a nearest neighbor in fragmentation space; and identifying a candidate for the parent ion using the nearest neighbor.[oni] Clause 23. A method of training a machine learning model for determining a molecular structure, including obtaining high dimensional mass spectra data for a plurality of known compounds, wherein, for each of the plurality of known compounds, the high dimensional mass spectra data is recorded at each collision energy of a series of collision energies and includes an abundance of one or more substructures of the known compound; composing, for each of the plurality of known compounds, a fragmentation profile using the high dimensional mass spectra data; embedding, for each of the plurality of known compounds, the high dimensional mass spectra data into a reduced dimension profile; and providing each reduced dimensionality fragmentation profile to the machine learning model for combining into a common fragmentation space.

[0112] Clause 24. The method of clause 23, wherein, once trained, the machine learning model infers the molecular structure of an unknown compound by fitting a fragmentation profile of the unknown compound into the common fragmentation space.

[0113] Clause 25. The method of clause 23 or 24, wherein the common fragmentation space has two dimensions.

[0114] Clause 26. The method of any one of clauses 23-25, wherein the fragmentation profile is configured as a heatmap.

[0115] Referring to the figures and examples presented herein generally, the disclosed environment provides a physical environment with which aspects of the molecular structural elucidation system can be implemented. Having described the preferred aspects and implementations of the present disclosure, modifications and equivalents of the disclosed concepts may readily occur to one skilled in the art. However, it is intended that such modifications and equivalents be included within the scope of the claims which are appended hereto.

Claims

What is claimed is:

1. A method of determining a molecular structure, the method comprising: obtaining mass spectrum data for an unknown parent ion, the mass spectrum data recorded at each collision energy of a series of collision energies and including an abundance of one or more substructures of the unknown parent ion; determining the molecular structure of the unknown parent ion by: composing the mass spectrum data into a fragmentation profde; embedding the fragmentation profde into a reduced dimension profde; mapping the reduced dimension profde into a fragmentation space; identifying one or more proximate molecules in the fragmentation space using the mapping; and identifying a candidate for the parent ion using the one or more proximate molecules.

2. The method of claim 1, wherein embedding the fragmentation profde into a reduced dimension profde comprises applying a Uniform Manifold Approximation & Projection (UMAP) algorithm to the fragmentation profde.

3. The method of claim 1 or 2, wherein determining the molecular structure of the parent ion is performed by a machine learning model.

4. The method of claim 3, wherein the machine learning model is trained by: obtaining high dimensional mass spectra data for a plurality of known compounds; embedding, for each of the plurality of known compounds, the high dimensional mass spectra data into a reduced dimension profde; providing each reduced dimension profde to the machine learning model for combining into the fragmentation space.

5. The method of claim 4, wherein the high dimensional mass spectra data comprises one or more fragmentation curves for substructures of the known compound.

6. The method of claim 5, wherein embedding the high dimensional mass spectra data into a reduced dimension profde comprises applying a Uniform Manifold Approximation & Projection (UMAP) algorithm to the heatmap.

7. The method of any one of claims 1-6, wherein the series of collision energies includes at least two collision energies, and preferably at least ten collision energies, and more preferably at least 16 collision energies.

8. The method of any one of claims 1-7, wherein the series of collision energies spans from 5 - 100 eV.

9. The method of any one of claims 1-8, wherein each of substructure of the one or more substructures is identified as being present at one or more collision energies of a series of collision energies, and preferably, at two or more collision energies of a series of collision energies, and more preferably at five or more collision energies of a series of collision energies.

10. The method of claim 9, wherein the collision energy, of the one or more collision energies, which has a highest intensity for the substructure is identified and added to the fragmentation profile.

11. The method of any one of claims 1-10, wherein the fragmentation profile comprises a dataframe spanning the mass range 50 - 550 Da.

12. The method of any one of claims 1-11, wherein the reduced dimension profile and the fragmentation space have a same number of dimensions, wherein the same number of dimensions is two dimensions.

13. The method of any one of claims 1-12, wherein embedding the fragmentation profile into a reduced dimension profile comprises applying one of a PCA and a tSNE algorithm to the fragmentation profile.

14. The method of any one of claims 1-13, wherein at least one of the one or more proximate molecules is a nearest neighbor and identifying a candidate for the parent ion using one or more proximate molecules comprises using the nearest neighbor.

15. A system for determining a molecular structure of a parent ion, the system comprising: a processor; a non-transitory memory in communication with the processor and storing instructions that, when executed, cause the processor to: obtain mass spectrum data for an unknown parent ion, the mass spectrum data recorded at each collision energy of a series of collision energies and including an abundance of one or more substructures of the unknown parent ion; determine the molecular structure of the unknown parent ion by: composing the mass spectrum data into a fragmentation profde; embedding the fragmentation profde into a reduced dimension profde; mapping the educed dimension profde into a fragmentation space and identifying a nearest neighbor in fragmentation space; and identifying a candidate for the parent ion using the nearest neighbor.

Citation Information

Cited By

  • Sewage high-risk pollutant screening and identification method based on fragmented tree pre-training

    CN121765490A

  • Method for screening and identifying high-risk pollutants in sewage based on fragmentation tree pre-training

    CN121765490B