Structure resolution based on machine learning and collision energy change

CN122603388APending Publication Date: 2026-08-18DH TECH DEVMENT PTE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580009916.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2025-01-14
Publication Date
2026-08-18

Smart Images

  • Figure CN122603388A_ABST
    Figure CN122603388A_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems and methods for determining a molecular structure, which includes obtaining mass spectrometry data of an unknown parent ion. The mass spectrometry data is recorded at each collision energy of a range of collision energies and includes abundances of one or more substructures of the unknown parent ion. The molecular structure of the unknown parent ion is determined by constructing the mass spectrometry data into a fragmentation map and embedding the fragmentation map into a reduced dimension map. The reduced dimension map is mapped into a fragmentation space. One or more proximal molecules are identified in the fragmentation space using the mapping and a parent ion candidate is identified using the one or more proximal molecules.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is filed as a PCT international patent application, claiming priority and benefit to U.S. Provisional Application No. 63 / 621,445, filed January 16, 2024, the disclosure of which is hereby incorporated by reference in its entirety. Background Technology

[0003] Mass spectrometry can be used to resolve molecular structures by providing valuable information about the mass and fragmentation patterns of ions originating from a given molecule. In structure resolution, the sample is first ionized, transforming it into charged particles. Depending on the properties of the compound, various ionization techniques are used, such as electrospray ionization (ESI) or matrix-assisted laser desorption / ionization (MALDI). The resulting ions are then accelerated by an electric or magnetic field, allowing them to be separated based on their mass-to-charge ratio (m / z). The obtained mass spectrum reflects the ion distribution, and the molecular ion peak provides the molecular weight of the compound.

[0004] Tandem mass spectrometry (MS / MS) can be used to reveal further details about molecular structure. By selecting specific ions from the mass spectrum and subjecting them to further fragmentation, the resulting fragment ions and their relative abundances provide insight into the atomic arrangement within the molecule. This fragmentation pattern helps identify functional groups, locate bond positions, and differentiate isomers. By integrating mass spectrometry with other spectroscopic techniques such as nuclear magnetic resonance (NMR) and infrared spectroscopy, researchers can gain a comprehensive understanding of the composition and structural characterization of compounds, facilitating their identification in complex mixtures or unknown samples. Summary of the Invention

[0005] The examples presented herein relate to a method for determining molecular structure. The method includes: obtaining mass spectrometry data of an unknown precursor ion, the mass spectrometry data being recorded at each of a series of collision energies and including the abundance of one or more substructures of the unknown precursor ion; determining the molecular structure of the unknown precursor ion by: constructing a fragmentation map from the mass spectrometry data; embedding the fragmentation map into a dimension-reduced map; mapping the dimension-reduced map into a fragmentation space; using the mapping to identify one or more neighboring molecules in the fragmentation space; and using the one or more neighboring molecules to identify precursor ion candidates.

[0006] In other aspects presented herein, embedding the fragmented graph into the dimensionality-reduced graph involves applying the Unified Manifold Approximation and Projection (UMAP) algorithm to the fragmented graph. In other aspects presented herein, the dimensionality-reduced graph includes a heatmap. In other aspects presented herein, the dimensionality-reduced graph has two dimensions.

[0007] In other aspects presented herein, the determination of the molecular structure of the parent ion is performed using a machine learning model. In yet another aspect presented herein, the machine learning model is trained by: obtaining high-dimensional mass spectrometry data of a variety of known compounds; for each of the known compounds, embedding the high-dimensional mass spectrometry data into a dimensionality-reduced graph; and providing each dimensionality-reduced graph to the machine learning model for combination into the fragmentation space. In still other aspects presented herein, the high-dimensional mass spectrometry data includes one or more fragmentation curves of the substructures of the known compounds. In still other aspects presented herein, embedding the high-dimensional mass spectrometry data into the dimensionality-reduced graph involves applying the Unified Manifold Approximation and Projection (UMAP) algorithm to the heatmap. In yet another aspect, the known compounds are small molecules.

[0008] In other aspects presented herein, the series of collision energies includes at least two collision energies, at least ten collision energies, or at least 16 collision energies. In still other aspects presented herein, the series of collision energies ranges from 5 to 100 eV.

[0009] In other aspects presented herein, each of the one or more substructures is identified as existing at one or more collision energies in a range of collision energies. In yet another aspect presented herein, each of the one or more substructures is identified as existing at two or more collision energies in a range of collision energies. In still some aspects presented herein, each of the one or more substructures is identified as existing at five or more collision energies in a range of collision energies. In still some aspects presented herein, the collision energy with the highest intensity for the substructure among the one or more collision energies is identified and added to the fragmentation pattern.

[0010] In other aspects presented herein, the fragmentation map comprises a data frame covering a mass range of 50–550 Da. In still other aspects presented herein, the reduced-dimensionality map and the fragmentation space have the same number of dimensions. In yet another aspect presented herein, the same number of dimensions is two dimensions.

[0011] In other aspects presented herein, embedding the fragmentation map into a dimensionality-reduced map involves applying one of the PCA and tSNE algorithms to the fragmentation map. In still other aspects presented herein, at least one of the one or more neighboring molecules is the nearest neighbor, and identifying precursor ion candidates using one or more neighboring molecules includes using the nearest neighbor.

[0012] Other examples presented herein relate to a system for determining the molecular structure of a precursor ion. The system includes a processor and a non-transitory memory, the non-transitory memory communicating with the processor and storing instructions that, when executed, cause the processor to: acquire mass spectrometry data of an unknown precursor ion, the mass spectrometry data being recorded at each of a series of collision energies and including the abundance of one or more substructures of the unknown precursor ion; determine the molecular structure of the unknown precursor ion by: constructing a fragmentation map from the mass spectrometry data; embedding the fragmentation map into a dimensionality-reduced map; mapping the dimensionality-reduced map into a fragmentation space and identifying nearest neighbors in the fragmentation space; and using the nearest neighbors to identify precursor ion candidates.

[0013] Other examples presented in this paper relate to a method for training a machine learning model to determine molecular structures. The method includes: obtaining high-dimensional mass spectrometry data of a plurality of known compounds, wherein for each of the plurality of known compounds, the high-dimensional mass spectrometry data is recorded at each of a series of collision energies and includes the abundance of one or more substructures of the known compounds; constructing a fragmentation map using the high-dimensional mass spectrometry data for each of the plurality of known compounds; embedding the high-dimensional mass spectrometry data into a dimensionality-reduced map for each of the plurality of known compounds; and providing each dimensionality-reduced fragmentation map to the machine learning model for combination into a common fragmentation space.

[0014] In other aspects presented herein, once trained, the machine learning model infers the molecular structure of the unknown compound by fitting the fragmentation map of the unknown compound to the common fragmentation space. In still other aspects presented herein, the common fragmentation space has two dimensions. In yet another aspect presented herein, the fragmentation map is configured as a heatmap. In other aspects, the known and unknown compounds are small molecules and do not include peptides or nucleic acid chains.

[0015] Several other aspects of the invention will also be set forth in the following description. These aspects may relate to individual features as well as combinations of features. It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory, and do not limit the broad inventive concept upon which the embodiments disclosed herein are based. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several aspects of this disclosure. A brief description of the drawings is as follows:

[0017] Figure 1 This is a block diagram of an example system for determining the structure of an unknown compound or identifying an unknown compound.

[0018] Figure 2 This is a block diagram of an example data processing system used to determine or identify molecular structures.

[0019] Figure 3 This is a flowchart of a method for determining molecular structure.

[0020] Figure 4 This is a flowchart of a method for determining the molecular structure of an unknown parent ion.

[0021] Figure 5 This is an example flowchart for generating stacked fragmentation curves based on mass spectrometry data.

[0022] Figure 6 yes Figure 4 Graphical illustrations of some aspects of the method.

[0023] Figure 7 It is an example of a fragmented pattern depicted as a heat map.

[0024] Figure 8 It is an instance of a dimensionality reduction map mapped to unknown objects in fragmented space.

[0025] Figure 9 This is a flowchart of a method used to train machine learning models to determine molecular structures.

[0026] Figure 10 Example block diagrams of virtual or physical computing systems are shown. Detailed Implementation

[0027] The way ions fragment in a mass spectrometer is largely determined by their molecular structure and the collision energy (CE) settings of the instrument. During the fragmentation process, molecules (e.g., small molecules) break down into their most stable units, many of which are substructures common to similar or even different molecular classes (e.g., benzene rings). The combination of substructures derived from a molecule is a fingerprint of that structure (or a group of very closely related structures) and can be analyzed manually or automatically to resolve the structure of the parent ion.

[0028] When performed manually by experts, determining the structure based on fragmentation patterns is time-consuming and incompatible with high-throughput work. However, no automated structure determination method has yet been developed that can match the accuracy achieved through manual fragmentation analysis. Compound identification using tandem mass spectrometry (MS / MS) and mass spectrometry is often extremely challenging when a perfect match does not exist in the existing spectral library. Users spend a significant amount of time on compound identification and typically require specialized knowledge. Even modest improvements to the current technology would be of considerable benefit, as it would enable less skilled personnel to obtain useful results in most cases.

[0029] The lack of accurate automated structure resolution stems from the challenges posed by the existence of defined substructures across different molecular classes. An additional dimension of mass spectrometry data is needed to resolve the true fragmentation tree from the mass spectrum. This paper discloses novel systems and methods for molecular identification based on characteristic fragmentation, providing faster and more efficient identification of the structure of the parent ion (and, in some cases, its identity) based on the fragmentation patterns generated by the parent ion.

[0030] Machine learning (ML) is well-suited for solving the problem of determining molecular structures from mass spectrometry because it excels at pattern sensing in high-dimensional, complex data. Unlike hard-coded algorithms (which are time-consuming and relatively inflexible), ML methods build upon a large number of parameters (e.g., 10^6). 3 -10 9 The model's parameters are adjusted by feeding training data into the model. In the case of mass spectrometry data, the training data typically consists of experimentally measured mass spectra paired with the identities of the molecules that generated the mass spectra. As an example, the model can learn from this to predict the structure of unknown molecules based on the spectra.

[0031] Careful design of training data is crucial for the development of ML models. As disclosed herein, spectra and molecules are represented in a novel way to present as much information as possible to the model. Individual spectra are acquired for each molecule under a large number of collisional CEs to provide in-depth insights into molecular fragmentation pathways. In this embodiment, data are acquired in a continuous slant mode. Spectra from different collisional CEs are not averaged but stacked into a 2D image representing information about collision-induced fragmentation of individual molecules.

[0032] The chemical space and molecular structures used to generate training data are most commonly represented as strings in informatics environments, such as Simplified Molecular Linear Input Specifications (SMILES) or International Compound Identifiers (InChI). However, these one-dimensional strings do not directly represent the three-dimensional structure of molecules and require ML models to learn their specific syntax. As disclosed in this paper, the molecules associated with each set of spectra are represented in multiple dimensions to allow the model to incorporate a higher degree of molecular structural information.

[0033] Dimensionality reduction is employed to generate manageable training datasets from large amounts of complex, high-dimensional data. Dimensionality reduction is a powerful tool for highlighting patterns in data that are too dimensional to be directly visualized. For example, a blood sample might be analyzed for 100 different biomarkers, and then reduced to a single dimension: the probability of having a specific disease. One algorithmic implementation of dimensionality reduction is the Unified Manifold Approximation and Projection (UMAP). UMAP uses, for example, 10... 2 -10 5Data in 2-10 dimensions, and embedding the data into a smaller number of dimensions. 2 The key point is that in the low-dimensional space, UMAP preserves both local and global structure. In other words, points that are close to each other in the embedding share the patterns detected by UMAP in the high-dimensional space.

[0034] While the advantages of this disclosure and the examples presented herein are described in relation to UMAP, other dimensionality reduction algorithms can be applied under the same principles. For example, one or more of the following may be suitable for dimensionality reduction in some implementations: t-distributed random nearest neighbor embedding (t-SNE), principal component analysis (PCA), locally linear embedding (LLE), isometric mapping (ISOMAP), autoencoders, thermal diffusion potential (PHATE) based on affinity-based trajectory embedding, diffusion mapping, or Laplacian eigenmaps. In embodiments, CE information may be simplified to a vector (e.g., increasing relative to decreasing) and used as input to ML.

[0035] UMAP and other dimensionality reduction algorithms are primarily used to reduce the dimensionality of data while preserving its underlying structure. It projects high-dimensional data points into a low-dimensional space (typically two or three dimensions) where the relationships between data points are preserved as much as possible. UMAP is also designed to capture the underlying manifold structure of the data. A manifold is a low-dimensional representation of complex high-dimensional data that captures the intrinsic relationships between data points. UMAP aims to find a representation that respects the topological structure and distances between points on the manifold. UMAP can reveal complex nonlinear patterns in data, making it applicable to a wider range of datasets.

[0036] As disclosed herein, high-dimensional data are generated from mass spectra of known compounds (e.g., small molecules) collected at different CEs. In embodiments, mass spectra may be collected at 2, 3, 4, 6, 10, 15, 16, 18, 20, etc., different CEs. In embodiments, different CEs may cover, for example, a range of 5-100 eV. For each compound, fragments are identified by looking for m / z values ​​appearing at multiple CEs. In embodiments, fragments may be identified by their appearance at at least two, at least three, at least five, at least ten, etc., different CEs. In a particular instance, fragments may be identified by their appearance at at least five different CEs.

[0037] For some or each of these fragments, the fragment with the highest intensity CE is identified and added to a data frame. In an embodiment, the data frame may cover a mass range of 50-550 Da at an appropriate m / z point. In instances using nominal masses, a 501-dimensional dataset is generated; however, those skilled in the art will understand that this can be extended to utilize high-resolution mass spectrometry data. An example table of a compound is shown in Table 1; entry 0 indicates that the compound does not form fragments at that mass:

[0038] Table 1

[0039]

[0040] According to this disclosure, a UMAP ML model is trained on 501-dimensional data, and the model learns to embed this data into a low-dimensional space (e.g., a two-dimensional space). The model is then applied to infer the structure of unknown compounds collected in the same high-dimensional format as the training data. The unknowns are embedded by the model into the same two-dimensional space. Inferences about the unknowns are based on distances and relationships to known compounds in the two-dimensional space. For example, nearest neighbors of the unknowns in the two-dimensional embedding are used to identify possible substructures within the unknowns. In embodiments, possible substructures can be identified using cheminformatics tools such as RDKIT.

[0041] The system and method disclosed in this paper offer numerous advantages over existing processes used for molecular identification employing multiple collision energies. For example, users can view visualizations of the degree of “similarity” in a range of three-dimensional structures of a molecule, determined based on the proximity of CE spectra of MS / MS fragments of the molecule. This multidimensional data representation reveals the relationships within the chemical space as a whole, capturing the relationships between compounds within that space. This representation provides more information and is superior to classic “unknown identification reports,” which typically identify numerous reference compounds in response to an unknown. Furthermore, the proposed system and method do not rely on library searches and can therefore be used to identify molecules that have not yet been cataloged. By correlating collision energies with substructure mass and then enabling the model to discover high-dimensional patterns, a wealth of mass information about molecular structures is incorporated into the analysis. Unknown compound structure resolution assists users in identifying novel metabolites in complex mixtures or in screening unknown substances in a high-throughput manner.

[0042] In particular, for embodiments incorporating UMAP, UMAP is typically faster than other dimensionality reduction algorithms, and on standard processing devices, embedding new spectra into existing models can be done in about a few seconds (e.g., about 2 seconds in some cases).

[0043] Figure 1This is a block diagram of an example system 100 for determining the structure or identifying unknown compounds. In embodiments, the system for determining the structure and identity of compounds may be integrated with a mass spectrometry system, as shown in example system 100, or in other embodiments, it may be independent of a mass spectrometry system. Example system 100 includes an ion source 110, a first mass separator 120, a fragmentation device 130, a second mass separator or mass analyzer 140, and a computing system 150.

[0044] In an embodiment, system 100 further includes a sample introduction device 160. The sample introduction device 160 introduces one or more compounds of interest from a sample into ion source 110 over time. The sample introduction device 160 performs techniques including, but not limited to, direct injection, liquid chromatography, gas chromatography, capillary electrophoresis, or ion mobility.

[0045] Mass filter 120 and fragmentation device 130 are shown as different stages of a quadrupole, and mass analyzer 140 is shown as a time-of-flight (TOF) device. Those skilled in the art will understand that either mass filter 120 or mass analyzer 140 may include other types of mass separators and analytical devices, including but not limited to ion traps, orbital traps, ion mobility devices, time-of-flight (TOF) devices, or Fourier transform ion cyclotron resonance (FT-ICR) devices. In embodiments, mass filter 120 and mass analyzer 140 are corresponding examples of a first mass separator and a second mass separator arranged in series. For example, a system may be configured according to this disclosure with a quadrupole for the first mass separator or mass filter 120 and a TOF device for the second mass separator or mass analyzer 140. Each mass separator is configured to receive a set of ions, detect the set of ions, and generate a set of detection signals corresponding to the detection of the set of ions.

[0046] Ion source device 110 converts a sample or a compound of interest from a sample into an ion beam. Ion source device 110 can perform ionization techniques including, but not limited to, matrix-assisted laser desorption / ionization (MALDI) or electrospray ionization (ESI).

[0047] Mass filter 120 receives the ion beam. In this embodiment, mass filter 120 is configured by the user for a specific precursor ion transport window based on the experimental objectives of the sample being analyzed. As discussed herein, a precursor ion transport window refers to the range of precursor or parent ions that are allowed to pass through a specific selection step and enter subsequent stages of mass analysis or fragmentation. In many tandem mass spectrometry (MS / MS) experiments, precursor ions are first selected based on their m / z (mass-to-charge ratio) to separate specific ions of interest for further analysis or fragmentation. The precursor ion selection process employs a mass filter or a specific set of voltages such that only ions within a certain m / z range (the precursor ion transport window) can pass through to the next stage.

[0048] The fragmentation device 130 of the tandem mass spectrometer 102 fragments or transports precursor ions transferred by the mass filter 120. In acquisition-related instances independent of data, this specifically includes scanning SWATH, generating one or more resulting product ions for each overlapping window of the series. When a collision energy sufficiently high to fragment the ions is used, the fragmentation device 130 fragments the precursor ions. When a collision energy sufficiently low to not fragment the ions is used, the fragmentation device 130 transports the precursor ions. Therefore, the resulting product ions may include precursor ions.

[0049] Collision energy influences ion fragmentation during collision-induced dissociation (CID) or collision-induced fragmentation (CIF). This process is commonly used in MS / MS to provide structural information about molecules by inducing the dissociation of parent or precursor ions into fragment ions. Collision energy is the kinetic energy gained by a precursor ion during a collision with a collision gas (such as helium or nitrogen) in the collision cell of a mass spectrometer.

[0050] Collision energies can be controlled by the instrument operator and can be adjusted to optimize ion fragmentation for specific analyses. Collision energies are typically expressed in electron volts (eV) or joules (J) and are specific to the instrument used and the type of collision cell. The collision energy setting affects the resulting fragmentation pattern and can influence the meaningfulness of any data acquired. For a single general experimental rule, too low a collision energy may result in insufficient fragmentation, while too high a collision energy may result in over-fragmentation, leading to the loss of important structural information.

[0051] By systematically increasing or decreasing collision energies, researchers can observe how fragmentation modes change. This method (often called collision energy skewing) allows for the identification of the optimal collision energies that produce the most informative and comprehensive structural information for a particular compound. Collision energy skewing can also be used to form fragmentation patterns of parent or precursor ions. By identifying precursor ions based on characteristic patterns across multiple collision energies (rather than a single collision energy CE), more structural data can be obtained.

[0052] The mass analyzer 140 of the tandem mass spectrometer 102 detects the intensity or count of each of one or more resulting product ions for each overlapping window of the series, the intensity or count forming mass spectrometry data for each overlapping window of the series.

[0053] The computing system 150 may be, but is not limited to, a computer, a microprocessor, Figure 7 The computing system 150 is a computing system or any device capable of sending and receiving control signals and data from and processing the data from the tandem mass spectrometer 102. The computing system 150 communicates with the ion source device 110, the mass filter 120, the fragmentation device 130, and the mass analyzer 140. The computing system 150 is shown as a separate device, but may be the processor or controller of the tandem mass spectrometer 102, or another device. The computing system 150 may store mass spectrometry data for each analyzed precursor ion window in a memory device (not shown), wherein, in an example performing a scan-type SWATH, mass spectrometry data for each overlapping window in the series is included. In an embodiment, the computing system 150 alternatively performs encoding and storage steps, encoding and storing each unique product ion detected by the mass analyzer 140 in real time during data acquisition. Before storing the mass spectrometry data, the computing system 150 performs one or more processing steps on the received raw mass spectrometry data to prepare the data for viewing, analysis, and storage. The raw mass spectrometry data includes the count or intensity of product ions at different m / z ratios over time.

[0054] Figure 2 This is a block diagram of an example data processing system 200 for determining or identifying molecular structures. In an embodiment, the data processing system 200 is implemented by a computing system 150. When operating the mass spectrometer 102 and processing the data output by the mass spectrometer, the data processing system 200 may be one of multiple subsystems or software executed by the computing system 150. This disclosure relates to determining molecular structures or identities; however, those skilled in the art will readily understand that other subsystems and / or modules may exist within and be executed by the computing system 150. The computing system 150 is presented as a single device in the examples herein, but in embodiments it may be one or more processing devices networked or otherwise communicating. Functionality may be partitioned between individual devices or shared across the collective processing capabilities of one or more processing devices. In an embodiment, the data processing system 200 forms part of or acts as a controller. The controller may be configured to issue operating commands to the mass spectrometer and / or receive unprocessed mass data including a set of detection signals from a second mass separator, and perform one or more data processing actions on the unprocessed mass data.

[0055] exist Figure 2In one example, the data processing system 200 includes a preprocessor 204, a dimensionality reduction 206, a fragmentation model 208, a fragmentation map 210, and a file formatter 212.

[0056] Raw mass spectrometer data are typically large and complex, requiring extensive data processing and analysis to extract meaningful information. Data processing can be performed as a series or set of actions. These actions can be performed collectively by the data processing system 200, or individual steps or portions of the processing can be performed by individual components or modules of the data processing system 200. In some embodiments, some components or features shown as integrated with the data processing system 200 may alternatively or additionally be performed at other locations on external components.

[0057] Preprocessor 204 performs one or more processing operations on mass spectrometry data 202 to prepare the data for indexing, storage, and subsequent analysis. Before indexing, mass spectrometry data typically undergoes preprocessing, which includes, for example, data transformation, noise reduction, peak extraction, and deconvolution. This step helps to simplify the data and improve the quality of the information to be indexed. Peaks are identified and quantified, representing ions and their corresponding intensities at specific m / z values. The detected peaks are typically used as the basis for indexing.

[0058] Dimensionality reduction 206 performs one or more dimensionality reduction algorithms on the mass spectrometry data 202. In an embodiment, dimensionality reduction 206 includes applying the UMAP algorithm to the high-dimensional mass spectrometry data. As discussed above, UMAP may provide some advantages in some embodiments, but other dimensionality reduction algorithms will also be applicable and suitable for embodiments of this disclosure. Other dimensionality reduction algorithms can be applied under the same principles. For example, one or more of t-distributed random neighbor embedding (t-SNE), principal component analysis (PCA), locally linear embedding (LLE), isometric mapping (ISOMAP), autoencoder, affinity-based trajectory embedding thermal diffusion potential (PHATE), diffusion mapping, or Laplacian eigenmaps may be suitable for dimensionality reduction in some implementations.

[0059] Fragmentation model 208 is a trained machine learning model used to analyze the fragmentation patterns of an unknown compound and map the compound into a fragmentation space. In an embodiment, fragmentation model 208 may incorporate aspects of dimensionality reduction 206.

[0060] The fragmentation model 208 can be trained by embedding high-dimensional mass spectrometry data into a reduced-dimensional spectrum for each of a plurality of known compounds. For each of the plurality of known compounds, the high-dimensional mass spectrometry data may include one or more fragmentation curves of the substructure of the known compound. For each of the plurality of known compounds, the mass spectrometry data may be collected at a range of collision energies. In embodiments, the range of collision energies includes at least two collision energies, at least ten collision energies, at least 16 collision energies, etc. In some cases, the range of the range of collision energies is 5-100 eV, although other suitable ranges of the range of collision energies are also contemplated, and those skilled in the art can apply such ranges according to experimental guidelines in various use cases.

[0061] Fragmentation space 210 is a stored fragmentation space generated by fragmentation model 208 and then used for inference of the analyzed unknown compounds. Fragmentation space 210 can be generated once and used for multiple inferences. In embodiments, fragmentation space 210 can be generated based on known compound profiles compiled for a specific experimental run or a series of experimental runs.

[0062] File formatter 212 converts the final processed mass spectrometry data into a standardized file format for storage and transmission. In this embodiment, the standardized file format is a Wiff file.

[0063] Figure 3 This is a flowchart of a method 300 for determining molecular structure. In an embodiment, method 300 is performed by a data processing system (such as...). Figure 2 The data processing system 200) is executed.

[0064] At operation 302, mass spectrometry data of the unknown precursor ion is obtained. The mass spectrometry data includes the abundance of one or more substructures of the unknown precursor ion. Mass spectrometry data typically provides information about the mass-to-charge ratio of the substructures of the resulting precursor ion, providing a fingerprint of the compound's molecular composition and structure through unique patterns of peaks and fragment ions. In embodiments, obtaining the mass spectrometry data involves receiving a processed data file from a mass spectrometer or storage system. In some cases, obtaining the mass spectrometry data includes collecting the data by a mass spectrometer.

[0065] The mass spectrometry data are recorded at each collision energy in a series of collision energies. In embodiments, the series of collision energies includes at least two collision energies, at least ten collision energies, at least 16 collision energies, etc. In some cases, the range of the series of collision energies is 5-100 eV, although other suitable ranges for the series of collision energies are also contemplated, and those skilled in the art can apply such ranges according to experimental guidelines for various use cases.

[0066] At operation 304, the molecular structure of the unknown precursor ion is determined. In this embodiment, determining the molecular structure of the unknown precursor ion may encompass identifying the precursor ion.

[0067] Figure 4 This is a flowchart of a method 400 for determining the molecular structure of an unknown precursor ion. Method 400 may be a sub-flow of step 304 of method 300 and represents the steps taken by the system to determine the molecular structure of the precursor ion. In embodiments, method 400 is performed by a machine learning model. In some cases, the machine learning model may be a sub-component of a data processing system that performs method 300, such as... Figure 2 The data processing system 200 has a fragmentation model 208.

[0068] The machine learning model can be trained by embedding high-dimensional mass spectrometry data into a reduced-dimensional spectrum for each of a plurality of known compounds. For each of the plurality of known compounds, the high-dimensional mass spectrometry data may include one or more fragmentation curves of the substructure of the known compound. For each of the plurality of known compounds, the mass spectrometry data may be collected at a range of collision energies. In embodiments, the range of collision energies includes at least two collision energies, at least ten collision energies, at least 16 collision energies, etc. In some cases, the range of the range of collision energies is 5-100 eV, although other suitable ranges of the range of collision energies are also contemplated, and those skilled in the art can apply such ranges according to experimental guidelines in various use cases.

[0069] In embodiments, embedding high-dimensional mass spectrometry data into a reduced-dimensional map includes applying the UMAP algorithm to the high-dimensional mass spectrometry data. As discussed above, UMAP may provide advantages in some embodiments, but other dimensionality reduction algorithms will also be applicable and suitable for embodiments of this disclosure. Each or a portion of the reduced-dimensional map is then provided to the machine learning model. The model then combines the reduced-dimensional maps into a fragmented space.

[0070] At operation 402, the mass spectrometry data is constructed into a fragmentation map. The high-dimensional mass spectrometry data contains one or more fragmentation curves representing the substructure of the unknown compound. The fragmentation curves represent the relationship between the intensity of ion fragments and the collision energy applied during collision-induced dissociation (CID) or other fragmentation processes. These curves illustrate how the abundance of a particular fragment ion varies with collision energy. According to this disclosure, the fragmentation curves can be stacked into a fragmentation map, rather than being averaged together or considered individually. In this way, multiple fragmentation curves from various CEs can be combined into a single overall map without losing the individual characteristics of each CE. Figure 5This is an example flowchart 500 for generating stacked fragmentation curves based on mass spectrometry data.

[0071] The fragmentation pattern can be generated as a data frame, as shown in the examples in Table 1 above. For example, in some embodiments, the fragmentation pattern includes a data frame covering a mass range of 50-550 Da.

[0072] In embodiments, each of the one or more substructures is identified as existing at one or more collision energies in a series of collision energies, at two or more collision energies in a series of collision energies, at five or more collision energies in a series of collision energies, and so on. For example, in some cases, each resulting substructure may be recorded in a fragmentation spectrum, or those substructures appearing at only one CE (or only two CEs, or fewer than five CEs) may be discarded to simplify the data and provide clearer results.

[0073] In an embodiment, the collision energy with the highest intensity for the substructure among the one or more collision energies is identified and added to the fragmentation spectrum. For example, once the substructure is identified as present at a sufficient number of collision energies (CEs), the substructure is recorded in the fragmentation spectrum at the CE with the highest intensity. This collision energy is populated into data frame cells corresponding to the compound and fragment mass. Examples are presented in Table 2 below.

[0074] Table 2

[0075]

[0076] In the examples in Table 2, “0” indicates that the compound does not produce fragments at a specific mass. For the mass at which the compound produced fragments, the CE (Crystal Scale) of maximum intensity is shown. For example, compound 1 produces a fragment with a mass of 120 Da, which has maximum intensity at 21 eV; and another fragment with a mass of 135 Da, which has maximum intensity at 8 eV. In some instances, only fragments between 50 and 550 Da are considered, but other mass ranges are considered and may be applicable given the experimental parameters for some applications of this disclosure. This example data frame or fragmentation plot is considered 501-dimensional because each entry is described by 501 variables.

[0077] Figure 6 yes Figure 4 The flowchart illustrates aspects of method 400. Multiple fragmentation curves 452 are combined to form a fragmentation map 454. In an embodiment, the fragmentation map 454 is depicted as a heatmap. In some cases, the reduced-dimensional map has two dimensions.

[0078] Figure 7This is an example of a fragmentation pattern depicted as a heatmap. As shown in this example, fragmentation curve data for a single molecule can be represented as a heatmap. In this example, the x-position represents the mass, where each vertical stripe is a fragment, and the y-position represents the collision energy. The color of each pixel can represent the normalized intensity of the measured mass / CE pair. The vertical stripes corresponding to each fragment may not be continuous because the data was not collected at every possible collision energy.

[0079] At operation 404, the fragmented graph is embedded into the reduced-dimensional graph. In an embodiment, embedding the fragmented graph into the reduced-dimensional graph includes applying the UMAP algorithm to the fragmented graph. In some cases, embedding the fragmented graph into the reduced-dimensional graph is achieved by applying one of the PCA and tSNE algorithms to the fragmented graph.

[0080] At operation 406, the reduced-dimensional map is mapped onto the fragmented space. From Figure 6 As can be seen, the ML model 456 receives the dimensionality-reduced graph and places it in the fragmentation space 458. In some preferred embodiments, the dimensionality-reduced graph and the fragmentation space have the same number of dimensions, which can be two dimensions.

[0081] At operation 408, mapping is used to identify one or more neighboring molecules in the fragmentation space. Figure 8 This is an example of a dimensionality-reduced map of an unknown object mapped into a fragmented space. In this example, unknown object 802 is mapped into fragmented space 804 using a UMAP model. The neighborhood region 806 surrounding unknown object 802 can be used to identify one or more neighbors 808. In an embodiment, at least one of the one or more neighboring molecules is the nearest neighbor, and identifying a parent ion candidate using one or more neighboring molecules involves using the nearest neighbor.

[0082] At operation 410, one or more neighboring molecules are used to identify candidate parent ions. Each of the one or more neighbors 808 has a known related structure, which the user or system can apply to infer the structure of the unknown 802. A trained expert or machine learning model can use the structures of neighboring known molecules to generate candidate structures for the unknown compound.

[0083] Figure 9 This is a flowchart of method 500 for training a machine learning model to determine molecular structure. At operation 502, high-dimensional mass spectrometry data for a variety of known compounds are obtained. For each known compound, high-dimensional mass spectrometry data is recorded at each collision energy in a series of collision energies. The mass spectrometry data typically includes the abundance of one or more substructures of the known compound.

[0084] At operation 504, for each of the known compounds, a fragmentation map is constructed using high-dimensional mass spectrometry data. In some cases, the fragmentation map is configured as a heatmap or data frame. At operation 506, the high-dimensional mass spectrometry data from each fragmentation map is embedded into the reduced-dimensional map.

[0085] At operation 508, each reduced-dimensional fragmentation map is provided to the machine learning model to be trained. As the model evaluates the received training data, patterns are identified, and each known compound is grouped into a common fragmentation space. In some cases, the common fragmentation space has two dimensions. This common fragmentation space can then be used to infer the structure and identity of unknown compounds evaluated using the model. In an embodiment, once trained, the machine learning model infers the molecular structure of the unknown compound by fitting the fragmentation map of the unknown compound to the common fragmentation space.

[0086] Figure 10 An example block diagram of a virtual or physical computing system 150 is shown. One or more aspects of the computing system 150 can be used to implement systems and methods for constructing data structures for high-dimensional extraction. Specifically, the computing system 150 can be used to implement a data processing system 200 and underlying or integrated components such as dimensionality reduction 206 and a fragmented model 208.

[0087] In the illustrated embodiment, the computing system 150 includes one or more processors 152, a system memory 158, and a system bus 172 coupling the system memory 158 to the one or more processors 152. The system memory 158 includes RAM (Random Access Memory) 160 and ROM (Read-Only Memory) 162. A basic input / output system containing basic routines is stored in the ROM 162, which facilitates, for example, the transfer of information between components within the computing system 150 during startup. The computing system 150 further includes a mass storage device 164. The mass storage device 164 is capable of storing software instructions and data. The mass storage device 164 may correspond to... Figure 2 The storage device 212, and as described in more detail above, stores index 214 and matrix 216. One or more processors 152 may be one or more central processing units or other processors.

[0088] Mass storage device 164 is connected to one or more processors 152 via a mass storage controller (not shown) connected to system bus 172. Mass storage device 164 and its associated computer-readable data storage media provide non-volatile, non-transitory storage for computing system 150. Although the description of computer-readable data storage media contained herein refers to mass storage devices such as hard disks or solid-state drives, those skilled in the art will understand that computer-readable data storage media can be any available non-transitory physical device or article of manufacture from which a central display station can read data and / or instructions.

[0089] Computer-readable data storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable software instructions, data structures, program modules or other data. Exemplary types of computer-readable data storage media include, but are not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Versatile Optical Disc), other optical storage media, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store the required information and can be accessed by the computing system 150.

[0090] According to various embodiments of the present invention, computing system 150 can operate in a networked environment using a logical connection to a remote network device via network 148. Network 148 is a computer network, such as an enterprise intranet and / or the Internet. Network 148 may include a LAN, a wide area network (WAN), the Internet, wireless transmission media, wired transmission media, other networks, and combinations thereof. Computing system 150 can be connected to network 148 via network interface unit 154 connected to system bus 172. It should be understood that network interface unit 154 can also be used to connect to other types of networks and remote computing systems. Computing system 150 also includes an input / output controller 156 for receiving and processing input from various other devices, including a touch user interface display or another type of input device. Similarly, input / output controller 156 can provide output to a touch user interface display or other types of output devices.

[0091] As briefly mentioned above, the mass storage device 164 and RAM 160 of the computing system 150 can store software instructions and data. The software instructions include an operating system 168 adapted to control the operation of the computing system 150. The mass storage device 164 and / or RAM 160 also store software instructions that, when executed by one or more processors 152, cause one or more of the systems, apparatuses, or components described herein to provide the functions described herein. For example, the mass storage device 164 and / or RAM 160 can store software instructions that, when executed by one or more processors 152, cause the computing system 150 to receive and execute processes for managing network access control and building systems.

[0092] The software instructions further include one or more software applications 166. Software applications may include dedicated systems and algorithms for performing specific tasks or actions or providing specific interfaces. Data processing system 200 and / or one or more components of data processing system 200 may be covered by software application 166.

[0093] Illustrative examples of the systems and methods described herein are provided below. Embodiments of the systems or methods described herein may include any one or more of the terms described below, and any combination thereof.

[0094] Clause 1. A method for determining a molecular structure, the method comprising: obtaining mass spectrometry data of an unknown precursor ion, the mass spectrometry data being recorded at each of a series of collision energies and including the abundance of one or more substructures of the unknown precursor ion; determining the molecular structure of the unknown precursor ion by: constructing a fragmentation map from the mass spectrometry data; embedding the fragmentation map into a dimension-reduced map; mapping the dimension-reduced map into a fragmentation space; identifying one or more neighboring molecules in the fragmentation space using the mapping; and identifying precursor ion candidates using the one or more neighboring molecules.

[0095] Clause 2. The method according to Clause 1, wherein embedding the fragmented map into the dimensionality-reduced map includes applying the Unified Manifold Approximation and Projection (UMAP) algorithm to the fragmented map.

[0096] Clause 3. The method according to Clause 2, wherein the dimensionality reduction map includes a heatmap.

[0097] Clause 4. The method described in Clause 2, wherein the dimensionality reduction graph has two dimensions.

[0098] Clause 5. The method according to any one of Clauses 1 to 4, wherein the determination of the molecular structure of the parent ion is performed by a machine learning model.

[0099] Clause 6. The method according to Clause 5, wherein the machine learning model is trained by: obtaining high-dimensional mass spectrometry data of a plurality of known compounds; for each of the plurality of known compounds, embedding the high-dimensional mass spectrometry data into a dimensionality-reduced spectrum; and providing each dimensionality-reduced spectrum to the machine learning model for combination into the fragmentation space.

[0100] Clause 7. The method according to Clause 6, wherein the high-dimensional mass spectrometry data includes one or more fragmentation curves of the substructure of the known compound.

[0101] Clause 8. The method according to Clause 6, wherein embedding the high-dimensional mass spectrometry data into a reduced-dimensional map includes applying the Unified Manifold Approximation and Projection (UMAP) algorithm to the heatmap.

[0102] Clause 9. The method according to any one of Clauses 1 to 8, wherein the series of collision energies includes at least two collision energies.

[0103] Clause 10. The method described in Clause 9, wherein the series of collision energies includes at least ten collision energies.

[0104] Clause 11. The method according to Clause 10, wherein the series of collision energies includes at least 16 collision energies.

[0105] Clause 12. The method according to any one of Clauses 1 to 11, wherein the range of collision energies is 5-100 eV.

[0106] Clause 13. The method according to any one of Clauses 1 to 12, wherein each of the one or more substructures is identified as existing at one or more collision energies in a series of collision energies.

[0107] Clause 14. The method according to Clause 13, wherein each of the one or more substructures is identified as existing at two or more collision energies in a series of collision energies.

[0108] Clause 15. The method according to Clause 14, wherein each of the one or more substructures is identified as existing at five or more collision energies in a series of collision energies.

[0109] Clause 16. The method according to Clause 13, wherein the collision energy with the highest intensity for the substructure is identified from the one or more collision energies and added to the fragmentation spectrum.

[0110] Clause 17. The method according to any one of Clauses 1 to 16, wherein the fragmentation spectrum includes a data frame covering a mass range of 50-550 Da.

[0111] Clause 18. The method according to any one of Clauses 1 to 17, wherein the dimensionality reduction graph and the fragmented space have the same number of dimensions.

[0112] Clause 19. The method described in Clause 18, wherein the same number of dimensions is two dimensions.

[0113] Clause 20. The method according to any one of Clauses 1 to 19, wherein embedding the fragmentation map into the dimensionality-reduced map comprises applying one of PCA and tSNE algorithms to the fragmentation map.

[0114] Clause 21. The method according to any one of Clauses 1 to 20, wherein at least one of the one or more neighboring molecules is the nearest neighbor, and the identification of a parent ion candidate using one or more neighboring molecules includes the use of the nearest neighbor.

[0115] Clause 22. A system for determining the molecular structure of a precursor ion, the system comprising: a processor; a non-transitory memory communicating with the processor and storing instructions, the instructions, when executed, causing the processor to: acquire mass spectrometry data of an unknown precursor ion, the mass spectrometry data being recorded at each of a series of collision energies and including the abundance of one or more substructures of the unknown precursor ion; determine the molecular structure of the unknown precursor ion by: constructing a fragmentation map from the mass spectrometry data; embedding the fragmentation map into a dimensionality-reduced map; mapping the dimensionality-reduced map into a fragmentation space and identifying nearest neighbors in the fragmentation space; and using the nearest neighbors to identify precursor ion candidates.

[0116] Clause 23. A method for training a machine learning model to determine molecular structures, the method comprising: obtaining high-dimensional mass spectrometry data of a plurality of known compounds, wherein for each of the plurality of known compounds, the high-dimensional mass spectrometry data is recorded at each of a series of collision energies and includes the abundance of one or more substructures of the known compounds; constructing a fragmentation map using the high-dimensional mass spectrometry data for each of the plurality of known compounds; embedding the high-dimensional mass spectrometry data into a dimensionality-reduced map for each of the plurality of known compounds; and providing each dimensionality-reduced fragmentation map to the machine learning model for combination into a common fragmentation space.

[0117] Clause 24. The method according to Clause 23, wherein once trained, the machine learning model infers the molecular structure of the unknown compound by fitting the fragmentation pattern of the unknown compound to the common fragmentation space.

[0118] Clause 25. The method described in Clause 23 or 24, wherein the common fragmentation space has two dimensions.

[0119] Clause 26. The method according to any one of Clauses 23 to 25, wherein the fragmentation pattern is configured as a heat map.

[0120] Generally, with reference to the accompanying drawings and examples presented herein, the disclosed environment provides a physical setting in which various aspects of a molecular structure resolution system can be implemented. After describing the preferred aspects and embodiments of this disclosure, modifications and equivalents to the disclosed concept will readily occur to those skilled in the art. However, such modifications and equivalents are intended to be included within the scope of the appended claims.

Claims

1. A method for determining molecular structure, the method comprising: Mass spectrometry data of an unknown precursor ion is obtained, wherein the mass spectrometry data is recorded at each of a series of collision energies and includes the abundance of one or more substructures of the unknown precursor ion. The molecular structure of the unknown parent ion was determined by the following: The mass spectrometry data is used to construct a fragmentation pattern; The fragmented map is embedded into the dimension-reduced map; The reduced-dimensional map is mapped onto the fragmented space; The mapping is used to identify one or more neighboring molecules in the fragmented space; as well as The parent ion candidates are identified using one or more neighboring molecules.

2. The method of claim 1, wherein embedding the fragmented map into the dimensionality-reduced map includes applying the Unified Manifold Approximation and Projection (UMAP) algorithm to the fragmented map.

3. The method according to claim 1 or 2, wherein determining the molecular structure of the parent ion is performed by a machine learning model.

4. The method of claim 3, wherein the machine learning model is trained by: Obtain high-dimensional mass spectrometry data of a variety of known compounds; For each of the known compounds, the high-dimensional mass spectrometry data is embedded into a reduced-dimensional spectrum; Each reduced-dimensional map is fed to the machine learning model to be combined into the fragmented space.

5. The method of claim 4, wherein the high-dimensional mass spectrometry data comprises one or more fragmentation curves of the substructure of the known compound.

6. The method of claim 5, wherein embedding the high-dimensional mass spectrometry data into the reduced-dimensional map includes applying the Unified Manifold Approximation and Projection (UMAP) algorithm to the heatmap.

7. The method according to any one of claims 1 to 6, wherein the series of collision energies includes at least two collision energies, and preferably at least ten collision energies, and more preferably at least 16 collision energies.

8. The method according to any one of claims 1 to 7, wherein the range of collision energies is 5-100 eV.

9. The method according to any one of claims 1 to 8, wherein each of the one or more substructures is identified as existing at one or more collision energies in a series of collision energies, and preferably at two or more collision energies in a series of collision energies, and more preferably at five or more collision energies in a series of collision energies.

10. The method of claim 9, wherein the collision energy with the highest intensity for the substructure among the one or more collision energies is identified and added to the fragmentation spectrum.

11. The method according to any one of claims 1 to 10, wherein the fragmentation spectrum comprises a data frame covering a mass range of 50-550 Da.

12. The method according to any one of claims 1 to 11, wherein the reduced-dimensional map and the fragmented space have the same number of dimensions, wherein the same number of dimensions is two dimensions.

13. The method according to any one of claims 1 to 12, wherein embedding the fragmentation map into the dimensionality-reduced map includes applying one of PCA and tSNE algorithms to the fragmentation map.

14. The method according to any one of claims 1 to 13, wherein at least one of the one or more neighboring molecules is the nearest neighbor, and identifying a parent ion candidate using one or more neighboring molecules includes using the nearest neighbor.

15. A system for determining the molecular structure of a parent ion, the system comprising: processor; A non-transitory memory, which communicates with the processor and stores instructions that, when executed, cause the processor to: Mass spectrometry data of an unknown precursor ion is obtained, wherein the mass spectrometry data is recorded at each of a series of collision energies and includes the abundance of one or more substructures of the unknown precursor ion. The molecular structure of the unknown parent ion was determined by the following: The mass spectrometry data is used to construct a fragmentation pattern; The fragmented map is embedded into the dimension-reduced map; The reduced-dimensional map is mapped into the fragmentation space, and the nearest neighbors are identified in the fragmentation space; as well as The nearest neighbor is used to identify parent ion candidates.