Predicting a mass spectrum from a chemical compound and uses thereof

EP4713958A2Pending Publication Date: 2026-03-25ENVEDA THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Current methodologies for predicting mass spectra of small molecules are computationally intensive and lack resolution, failing to accurately predict intensity values, and often rely on limited or inaccurate fragmentation techniques.

Method used

The use of machine learning models, specifically graph neural networks, to generate mass spectrum fragments from chemical structures by encoding molecular graphs and predicting mass-to-charge values and intensities, leveraging a limited vocabulary of compound fragments and neutral loss fragments.

Benefits of technology

This approach enables accurate and efficient prediction of mass spectra, improving computational efficiency and resolving the limitations of existing methods by generating high-resolution predictions of mass spectra for small molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024030098_28112024_PF_FP_ABST
    Figure US2024030098_28112024_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure, in certain aspects, is directed to methods of generating a mass spectrum from a chemical structure of a compound using a limited vocabulary of fragments and a machine learning model. Further provided, in some examples, are methods using the predicted spectrum or spectra, such as for generating a spectral library and / or identifying a compound associated with a mass spectrum.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No.: 226922001340 PREDICTING A MASS SPECTRUM FROM A CHEMICAL COMPOUND AND USES THEREOF CROSS-REFERENCE TO RELATED APPLICATOINS

[0001] This application claims priority to, and the benefit of, U.S. provisional application 63 / 467,856, filed on May 19, 2023, the contents of which are incorporated by reference in their entirety. TECHNICAL FIELD

[0002] The present disclosure, in certain aspects, is directed to methods of generating a mass spectrum from a chemical structure of a compound using a limited vocabulary of mass spectrum fragments and a machine learning model. Further provided, in some examples, are methods for predicting a mass spectrum of the compound and using the predicted mass spectrum for generating / updating a spectral library and / or identifying a compound associated with a mass spectrum. BACKGROUND

[0003] Tandem mass spectrometry is useful for the identification and quantification of small molecules and various aspects associated with properties in experiments. The general workflow for tandem mass spectrometry of a small molecule is (a) weigh the small molecule to obtain a mass; (b) fragment the small molecule; and (c) weigh the resulting fragments of the small molecule to obtain masses of each fragment. Measurements of the small molecule, or fragments thereof, include the mass-to-charge (m / z) value of an ion of the small molecule, or ions of fragments thereof, and the intensity or abundance of each ion. Such measurements can be reported or represented in a spectrum of intensity versus m / z.

[0004] It can be useful to perform an in silico prediction of a mass spectrum of a small molecule based on the structure of the small molecule. Current methodologies for mass spectrum prediction is computationally intensive (e.g., fragmentation trees) and / or lacks the desired resolution (e.g., rounding-based techniques). Moreover, certain methodologies only provides predicted m / z values, and cannot accurately predicted associated intensity values. 1sf-5954611Docket No.: 226922001340 BRIEF SUMMARY

[0005] Described herein are methods, systems, programming, and other techniques, for generating mass spectrum fragments based on a chemical structure of a compound. In some examples, first graph data representing a first chemical structure of a first compound may be obtained. Using one or more machine learning models, a first encoded representation of the first chemical structure may be generated based the first graph data, wherein the first encoded representation comprises a first plurality of graph features. Using the one or more machine learning models, a first plurality of mass spectrum fragments may be generated at least based on (i) the first encoded representation of the first chemical structure of the first compound and (ii) a first identified set of compound fragments. The first identified set of compound fragments may be stored as a vocabulary associated with the one or more machine learning models. The first identified set of compound fragments may at least be based on a predetermined set of compound fragments comprising first charged compound fragments and a first derived set of neutral loss compound fragments. The first derived set of neutral loss compound fragments may be derived at least based on a mass of one or more neutral fragments of the predetermined set of compound fragments and a mass of the first chemical structure of the first compound, and each first derived neutral loss compound fragment is neutral.

[0006] In some embodiments, generating the first plurality of mass spectrum fragments comprises: generating, using the one or more machine learning models, a plurality of mass-to- charge values and a plurality of intensities, wherein each mass spectrum fragment comprises one or more of the plurality of mass-to-charge values and one or more of the plurality of intensities. In some embodiments, the method further comprises generating, using the one or more machine learning models, a prediction of a mass spectrum of the first compound at least based on the first plurality of mass spectrum fragments. In some embodiments, the method further comprises identifying a set of most-probable compound fragments from the first chemical structure, wherein the predetermined set of compound fragments comprises the set of most-probable compound fragments.

[0007] In some embodiments, the method further comprises selecting the predetermined set of compound fragments from a historical dataset of experimentally observed compound fragments associated with the first compound. In some embodiments, selecting comprises: determining a number of observations associated with each of the experimentally observed compound fragmentssf-5954611Docket No.: 226922001340 from the historical dataset; generating a ranking of the experimentally observed compound fragments at least based on the number of observations associated with each of the experimentally observed compound fragments, wherein one or more of the experimentally observed compound fragments are selected from the historical dataset for the predetermined set of compound fragments at least based on the number of observations of each of the one or more of the experimentally observed compound fragments being greater than or equal to a threshold number of observations. In some embodiments, the predetermined set of compound fragments comprises 15,000 or less compound fragments, 10,000 or less compound fragments, 5,000 or less compound fragments, or 1,000 or less compound fragments.

[0008] In some embodiments, the method further comprises generating the first identified set of compound fragments by: selecting a set of compound fragments from a historical dataset of experimentally observed compound fragments associated with the compound; storing the selected set of compound fragments to the vocabulary; and updating the vocabulary at least based on the first derived set of neutral loss compound fragments, wherein the first identified set of compound fragments comprises the updated vocabulary. In some embodiments, the method further comprises generating the first derived set of neutral loss compound fragments by: subtracting a mass of one or more neutral loss compound fragments of the predetermined set of compound fragments from the mass of the first chemical structure. In some embodiments, the one or more neutral loss compound fragments comprise a top-K set of neutral fragments from the mass of the first chemical structure. In some embodiments, K equals 10,000 or less neutral fragments, 5,000 or less neutral fragments, or 1,000 or less neutral fragments.

[0009] In some embodiments, the one or more machine learning models comprise an encoder. In some embodiments, the encoder comprises at least one of: a graph neural network (GNN), a graph convolutional network (GCN), or a graph isomorphism network (GIN). In some embodiments, the one or more machine learning models comprise a decoder. In some embodiments, the decoder can comprise at least one of a transformer-based decoder, a convolutional neural network (CNN) decoder, or a feed-forward network decoder.

[0010] In some embodiments, the method further comprises generating the first graph data representing the chemical structure at least based on one or more simplified molecular-input line- entry system (SMILES) strings corresponding to the compound. In some embodiments, the method further comprises generating the first graph data representing the chemical structure atsf-5954611Docket No.: 226922001340 least based on one or more International Chemical Identifier (InChi) strings corresponding to the compound.

[0011] In some embodiments, the method further comprises obtaining second graph data representing a second chemical structure of a second compound different than the first compound; and updating the vocabulary based on a second derived set of neutral loss compound fragments, wherein the second derived set of neutral loss compound fragments is at least based on the mass of the one or more neutral fragments of the predetermined set of compound fragments and a mass of the second chemical structure. In some embodiments, the second derived set of neutral loss compound fragments is generated by subtracting the mass of the one or more neutral fragments of the predetermined set of compound fragments from the mass of the second chemical structure. In some embodiments, the method further comprises generating, using the one or more machine learning models, a second encoded representation of the second chemical structure of the second compound at least based on the second graph data, wherein the second encoded representation of the second chemical structure comprises a second plurality of graph features; and generating, using the one or more machine learning models, a second plurality of mass spectrum fragments at least based on (i) the second encoded representation of the second chemical structure and an identified second set of compound fragments, wherein the identified second set of compound fragments is at least based on the predetermined set of compound fragments and the second derived set of neutral loss compound fragments.

[0012] In some embodiments, the method further comprises obtaining the compound. In some embodiments, the compound is a natural product or a derivative thereof. In some embodiments, the compound is of a plant extract or a derivative of the compound. In some embodiments, the compound is a small molecule having a molecular weight of less than 2,000 Dalton (da). In some embodiments, the first graph data comprises information associated with one or more experimental covariates. In some embodiments, the one or more experimental covariates comprises information of a type of mass spectrometer used to generate training data for training the one or more machine learning models. In some embodiments, the one or more experimental covariates comprises information of a setting of a mass spectrometer used to generate training data for training the one or more machine learning models. In some embodiments, the setting of the mass spectrometer includes one or more of a collision energy, a collision gas, or an ionization type.sf-5954611Docket No.: 226922001340

[0013] Disclosed herein are additionally methods of generating a spectral library, the method comprising: performing the method of any one of the methods described herein on a plurality of compounds to generate one or more mass spectra associated with each of the plurality of compounds; and compiling the one or more mass spectra of each of the plurality of compounds to generate the spectral library. In some embodiments, the plurality of compounds comprises a known compound. In some embodiments, the plurality of compounds comprises a theoretical compound.

[0014] Disclosed herein are additionally methods of identifying a compound associated with a mass spectrum, the method comprising: analyzing the mass spectrum using the spectral library of according to any one of the methods described herein to identify the compound associated with the mass spectrum. In some embodiments, the method further comprises performing a mass spectrometry technique to obtain the mass spectrum.

[0015] Disclosed herein are additionally methods of identifying a compound associated with a mass spectrum, the method comprising: analyzing the mass spectrum using the spectral library according to any of the embodiments herein, to train a machine learning model to a predict a structure of the compound.

[0016] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods and / or part or all of one or more processes disclosed herein.

[0017] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosedsf-5954611Docket No.: 226922001340 can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.

[0018] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0020] FIG. 1 shows an example system for generating mass spectrum fragments, in accordance with various embodiments.

[0021] FIG.2 shows an example workflow of tandem mass spectrometry, in accordance with various embodiments.

[0022] FIGS. 3A-3B show an example chemical structure and mass spectrum, respectively, in accordance with various embodiments.

[0023] FIG. 4 shows a plot of signals of small molecule spectra as a function of vocabulary size, in accordance with various embodiments.

[0024] FIG. 5 shows an example graph neural network for predicting mass spectrum fragments, in accordance with various embodiments.

[0025] FIG.6 shows an example process for training a machine learning model, in accordance with various embodiments.

[0026] FIG. 7 shows example plots of a predicted mass spectra and an experimentally determined mass spectra for various example molecules, in accordance with various embodiments.

[0027] FIG. 8 shows an example flowchart of a method for generating mass spectrum fragments, in accordance with various embodiments.sf-5954611Docket No.: 226922001340

[0028] FIG. 9 shows an example mass spectra query process, in accordance with various elements.

[0029] FIG. 10 shows an example computing system used to implement one or more of the embodiments described herein. DETAILED DESCRIPTION

[0030] Provided herein, in certain aspects, are methods, programming, and systems for generating mass spectrum fragments from a chemical structure using a limited vocabulary of mass spectrum fragments and one or more machine learning models. The machine learning models may be trained to generate mass spectrum fragments based on the chemical structure and the limited vocabulary. The disclosure of the present application is based on the inventors’ unique perspectives and findings regarding methods for accurately predicting m / z values and associated intensities of tandem mass spectrometry fragments associated with a chemical structure, such as a small therapeutic compound, including as a function of any combination of experimental parameters, e.g., instrumentation and / or collision energy. The methodologies developed and taught herein uses a vocabulary composed of a limited set of possible fragments identified as the most common fragments and / or fragments of interest obtained from observed experimental fragmentation spectra. Such vocabulary can be numerically limited, e.g., 10,000 fragments (i.e., chemical formulas), and / or based on context, e.g., fragments relevant to human metabolism or an industrial setting. Using this vocabulary, the methods taught herein use trained machine models to predict possible fragmentation mass spectrum for an input chemical structure. Following this technique, many architectures are possible for executing the described workflows. For example, a dynamic architecture is envisioned where a starting vocabulary is used to generate a compound- specific vocabulary that is composed of charged fragments of the starting vocabulary that are present based on the chemical structure of the compound and any charged fragments of the compound that would result from a neutral loss present in the starting vocabulary, i.e., the vocabulary provided to the machine model is based on the input chemical structure. Use of the taught vocabulary-based techniques enables a high accuracy and speed of prediction while also improving computational efficiency. Downstream uses of predicted spectra can be appreciated by those of ordinary skill in the art, and such uses are encompassed in the description provided herein. For example, spectral libraries containing a predicted spectrum or predicted spectra, are useful for small molecule compound identification. Accurate in silico prediction of a tandem mass spectrumsf-5954611Docket No.: 226922001340 associated with a chemical structure can provide significant technical benefits, such as bypassing time, decreasing costs, and reducing effort in intensive experimental procedures needed for experimentally determining a tandem mass spectrum (which may include isolating enough starting material of a compound to generate an experimental tandem mass spectrum).

[0031] Thus, provided herein, in certain aspects, is a method for generating mass spectrum fragments based on a chemical structure of a compound. The method may comprise obtaining first graph data representing a first chemical structure of a first compound; generating, using one or more machine learning models, a first encoded representation of the first chemical structure based the first graph data, wherein the first encoded representation comprises a first plurality of graph features; and generating, using the one or more machine learning models, a first plurality of mass spectrum fragments at least based on (i) the first encoded representation of the first chemical structure of the first compound and (ii) a first identified set of compound fragments, wherein: the first identified set of compound fragments is stored as a vocabulary associated with the one or more machine learning models, the first identified set of compound fragments is at least based on a predetermined set of compound fragments comprising first charged compound fragments and a first derived set of neutral loss compound fragments, the first derived set of neutral loss compound fragments is derived at least based on a mass of one or more neutral fragments of the predetermined set of compound fragments and a mass of the first chemical structure of the first compound, and each first derived neutral loss compound fragment is neutral. I. Definitions

[0032] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.

[0033] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope ofsf-5954611Docket No.: 226922001340 the claimed subject matter. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For instance, where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, unless the context clearly dictate otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure. In some embodiments, two opposing and open ended ranges are provided for a feature, and in such description, it is envisioned that combinations of those two ranges are provided herein. For example, in some embodiments, it is described that a feature is greater than about 10 units, and it is described (such as in another sentence) that the feature is less than about 20 units, and thus, the range of about 10 units to about 20 units is described herein.

[0034] The term “about” as used herein refers to the usual error range for the respective value readily known in this technical field. Reference to “about” a value or parameter herein includes (and describes) variations that are directed to that value or parameter per se. For example, description referring to “about X” includes description of “X.”

[0035] As used herein, including in the appended claims, the singular forms “a,” “or,” and “the” include plural referents unless the context clearly dictates otherwise. For example, “a” or “an” means “at least one” or “one or more.” It is understood that aspects and variations described herein include embodiments “consisting” and / or “consisting essentially of” such aspects and variations.

[0036] As used herein, a “subject” or an “individual,” which are terms that are used interchangeably, is a mammal. In some embodiments, a “mammal” includes humans, non-human primates, domestic and farm animals, and zoo, sports, or pet animals, such as dogs, horses, rabbits, cattle, pigs, hamsters, gerbils, mice, ferrets, rats, cats, monkeys, etc. In some embodiments, the subject or individual is human.

[0037] Those skilled in the art will recognize that several embodiments are possible within the scope and spirit of the present disclosure. The following description illustrates the disclosure and,sf-5954611Docket No.: 226922001340 of course, should not be construed in any way as limiting the scope of the inventions described herein. II. Methods for generating mass spectrum fragments based on a chemical structure of a compound

[0038] In certain aspects, provided herein is a method for generating a mass spectrum from a chemical structure of a compound. Identifying unknown small molecules in complex chemical mixtures is an ongoing challenge for many applications, such as metabolomics, drug discovery, clinical diagnostics, forensics, environmental monitoring, and others. Small molecule identification is traditionally accomplished with the aid of a tandem mass spectrometry (also referred to as MS / MS). However, a key bottleneck in tandem mass spectrometry is structural elucidation: given a mass spectrum, determining the 2D structure of the molecule it represents. Typically, only 2-4% of spectra are identified in untargeted metabolomics experiments, and a recent competition saw no more than 30% accuracy.

[0039] Complicating matters are the facts that tandem mass spectrometry is a lossy measurement and existing training sets are small. This makes it challenging to form direct prediction of structures from mass spectra is particularly challenging.

[0040] A common approach to small molecule identification is spectral library search. This further morphs the small molecule identification problem into information retrieval task: given an observed spectrum, query the observed spectrum against a library of spectra associated with known structures.

[0041] One downside to this approach is that there are relatively few small molecules (of the order of 104small molecules) with known experimental mass spectra. Therefore, in spectral library searches, it is common to supplement the library with augment libraries of spectra predicted from large databases (106- 109) of molecular graphs. Therefore, one key aspect of the small molecule identification task is the process of predicting mass spectra, referred to as “spectrum prediction,” to be used as an augment library.

[0042] Spectrum prediction is actively studied in metabolomics and quantum chemistry, however, is fairly nascent to the field of machine learning. One of the challenges in spectrum prediction is the modelling of the output space. A mass spectrum is a variable-length set of real-sf-5954611Docket No.: 226922001340 valued (m / z, height) tuples, which is not straightforward to represent as an output of a deep learning model. The m / z coordinate (i.e., the “mass-to-charge” ratio) poses particular difficulty, as the m / z coordinate needs to be predicted with high precision. However, one of the benefits of tandem mass spectrometry is that it allows for fractional m / z differences to be identified on the order of 10-6, thereby distinguishing different elemental compositions.

[0043] Existing approaches – mass-binning, bond-breaking – to spectrum prediction force a tradeoff between capturing high-resolution m / z information and tractability of the learning problem. In an example, mass-binning represents a mass spectrum as a fixed-length vector by discretizing the m / z axis at regular intervals (e.g., binning), which can discard valuable information in favor of tractable learning. As another example, bond-breaking achieve perfect m / z resolution, but does so through expensive combinatorial enumeration of substructures.

[0044] Described herein are systems, methods, and programming for formulating spectrum prediction as a mapping from a molecular graph to a probability distribution over molecular formulas. Doing so can allow full resolution predictions to be made without enumerating substructures. One of the key discoveries of the described techniques is that most mass spectra can be effectively approximated with a small, fixed vocabulary of molecular formulas. This realization enables the tradeoff between m / z resolution and tractable learning to be bypassed.

[0045] In some embodiments, a molecular graph G, may be expressed mathematically as G = (V, E, a, b). The molecular graph G may represent a minimal description of the structure of a molecule. Molecular graph G may include a set of nodes V, a set of edges E, node labels a (where a ^

[0118] V), and edge labels e (where b ^ ({1, 1.5, 2, 3} × {-1, 0, 1})Eindicating bond order and chirality).

[0046] In some embodiments, graph data may be generated that represents molecular graph G. In one or more examples, the graph data may include data representing nodes V, edges E, node labels a, and edge labels e. In one or more examples, the graph data may include an adjacency matrix. The adjacency matrix may be a V × V matrix having values aij where aij = 1 if there is an edge E between node i and node j. In one or more examples, aij= w where w equals the weight of the edge E between node i and node j. 11sf-5954611Docket No.: 226922001340

[0047] In some embodiments, a molecular formula f (e.g., C8H10N4O2) may be used to express a multiset of atoms. Formulas may be added and subtracted from one another, and inequalities between formulas are taken to hold elementwise. In some embodiments, a chemical structure of a compound may be represented via a molecular formula. In one or more examples, the terms “molecular formula” and “chemical formula” may be used interchangeably.

[0048] In some embodiments, a precursor ion formula P refers to a molecule to be fragmented. The sub-formulas of P may form the set F(P). A mass spectrum is implicitly always accompanied by precursor formula P. The ions produces as a result of the tandem mass spectrometry process can be referred to herein as product ions.

[0049] In some embodiments, the theoretical mass of a molecule with formula f, with units of Daltons (Da), may be represented as a linear combination of the monoisotopic masses of the elements of the periodic table, with multiplicities given by f.

[0050] In some embodiments, a mass spectrum S is a variable-length set of peaks, each of which is a (m / z, height) tuple.

[0051] Furthermore, in one or more examples, unless otherwise indicated, as defined herein, charge z is set to be equal to 1, as it is rare for small molecules to acquire more than a single charge.

[0052] In some embodiments, the chemical structure of a compound may include its chemical formula.

[0053] FIG. 1 shows an example system 100 for generating mass spectrum fragments, in accordance with various embodiments. System 100 may include a computing system 102, a mass spectrometer 120, client devices 130-1 to 130-N (also referred to collectively as “client devices 130” and individually as “client device 130”), databases 140 (e.g., mass spectra database 142, training data database 144, model database 146, historical data database 148), or other components. In some embodiments, components of system 100 may communicate with one another using network 150, such as the Internet.

[0054] Client devices 130 may communicate with one or more components of system 100 via network 150 and / or via a direct connection. Client devices 130 may be a computing device configured to interface with various components of system 100 to control one or more tasks, cause one or more actions to be performed, or effectuate other operations. For example, client devicesf-5954611Docket No.: 226922001340 130 may be configured to input a query mass spectrum to computing system 102 to obtain results indicating a top-k most-likely matching mass spectra. Example computing devices that client devices 130 may correspond to include, but are not limited to, which is not to imply that other listings are limiting, desktop computers, servers, mobile computers, smart devices, wearable devices, cloud computing platforms, or other client devices. In some embodiments, each client device 130 may include one or more processors, memory, communications components, display components, audio capture / output devices, image capture components, or other components, or combinations thereof. Each client device 130 may include any type of wearable device, mobile terminal, fixed terminal, or other device.

[0055] It should be noted that while one or more operations are described herein as being performed by particular components of computing system 102, those operations may, in some embodiments, be performed by other components of computing system 102 or other components of system 100. As an example, while one or more operations are described herein as being performed by components of computing system 102, those operations may, in some embodiments, be performed by aspects of client devices 130. It should also be noted that, although some embodiments are described herein with respect to machine learning models, other prediction models (e.g., statistical models or other analytics models) may be used in lieu of or in addition to machine learning models (e.g., a statistical model replacing a machine learning model and a non- statistical model replacing a non-machine learning model in one or more embodiments). Still further, although a single instance of computing system 102 is depicted within system 100, additional instances of computing system 102 may be included (e.g., computing system 102 may comprise a distributed computing system).

[0056] Mass spectrometer 120 may be configured to produce a mass spectrum of a compound. In some embodiments, mass spectrometer 120 is a tandem mass spectrometer. Although only a single mass spectrometer 120 is included in FIG. 1, persons of ordinary skill in the art will recognize that system 100 may include additional mass spectrometers.

[0057] Mass spectrometer 120 may be cable of performing tandem mass spectroscopy to an input compound. As an example, with reference to FIG.2, tandem mass spectroscopy process 200 take an input compound 202 and produce a histogram 216 representing a mass spectrum of compound 202. In some embodiments, mass spectrometer 120 may be configured to ionize a chemical sample, compound 202, to produce a jet of electrically-charged gas 204a-204c. Gassf-5954611Docket No.: 226922001340 204a-204c may be electromagnetically filtered using electromagnetic components 206, such as a quadrupole, to select a set of precursor ions 208a-208c having a mass-to-charge ratio (m / z), each representing a particular molecular structure. Each precursor ion 208a-208c can be fragmented by, e.g., colliding precursor ion 208a-208c with molecules of an inert gas. If a collision occurs with sufficient energy, one or more bonds in precursor ion 208a-208c can break, producing a charged product ion 210a, 212a, 214b and, in some embodiments, one or more uncharged neutral loss molecules 210b, 212b, 214b. Charged product ions 210a, 212a, 214a can be measured by one or more detectors of mass spectrometer 120 as the detectors can record a measurement of mass- to-charge ratio (m / z) of product ions 210a, 212a, 214a (up to a small measurement error proportional to the mass-to-charge ratio (m / z) times the resolution į. This process may be repeated for large quantities of identical precursor ions 208a-208c to build a histogram 216 indexed by mass-to-charge ratio (m / z). Local maxima 218a-218c in histogram 216 may represent a unique product ion, with height reflecting its probability of formation. For example, product ion 210a may produce local maxima 218a, and the height of local maxima 218a may reflect the probability of product ion 210a being formed. Similarly, product ions 212a and 214a may produce local maxima 218b and 218c, respectively. This set of peaks, local maxima 218a-218c constitutes the mass spectrum of compound 202. A typical mass spectrometry experiment acquires mass spectra for tens of thousands of distinct precursors in this manner, with mass-to-charge ratios (m / z) measured at a resolution on the order of 10-6.

[0058] FIGS. 3A-3B show an example precursor ion 300 and mass spectrum 350, respectively, in accordance with various embodiments. In the example of FIG.3A, precursor ion 300 may have the formula: C8H11N4O2. As a result of the fragmentation process, some of the bonds between atoms / molecules within the precursor ion may be break. In other words, the molecular graph is “cut” into different components: one retaining charge and the other being neutral. For example, with reference to FIG. 2, precursor ion 208a may be fragmented into a product ion 210a and a neutral loss 210b. The bonds that break, and thus the ions born from the fragmentation, may be a result of the strength of the bonds between certain atoms of the precursor ion.

[0059] With reference back to FIG. 3A, fragmentation 302 of precursor ion 300 (e.g., C8H11N4O2) may produce a product ion 304 (e.g., C3H4NO2+) and a neutral loss 306 (e.g., C5H7N3). The ions that are produced provide insight into the chemical structure of the precursorsf-5954611Docket No.: 226922001340 ion. For example, because fragmentation 302 broke a bond between a first atom and a second atom, this indicates that the bond between the first atom and the second atom is weaker than other bonds of the precursor ion.

[0060] Looking at mass spectrum 350 of FIG. 3B, it can be seen that there is a first peak 352 at (m / z) = m and a second peak 354 at (m / z) = M. Each of peaks 352-354 may correspond to a mass value. Thus, which of peaks 352 and 354 correspond to product ion 304 and precursor ion 300 depends on the (m / z) value of each peak and the individual masses of each element forming precursor ion 300 of FIG. 3A. For example, first peak 352 may correspond to product ion 304 because the mass of each element forming product ion 304, when summed, equals (m / z) = m. Second peak 354 may represent precursor ion 300 (e.g., (m / z) = M), as the fragmentation process may leave some portion of the original compound. Second peak 354 can be mapped to product ion 304 or precursor ion 300 by summing the mass of each element of precursor ion 300 (e.g., 8(mC) 11(mH) 4 (mN) 2(mO) = M).

[0061] In addition to providing insight into the product ion produced by the fragmentation process, mass spectrum 350 may also indicate a neutral loss associated with the produced product ion. As detailed above with respect to FIG. 2, the fragmentation process of tandem mass spectroscopy produces a charged product ion and a neutral loss. In the example of FIGS.3A-3B, precursor ion 300 is fragmented to produce product ion 304 and neutral loss 306. The mass of neutral loss 306 may be determined based on a difference between M – the mass of precursor ion 300 – and m – the mass of product ion 304. In some examples, the fragmentation process may also generate product ions having an (m / z) =– m. In this scenario, a peak would exist at (m / z) = M – m instead of at m. In this example, the neutral loss would then have a mass m.

[0062] Computing system 102 may include a graph data generation subsystem 110, a mass spectrum fragment determination subsystem 112, a model training subsystem 114, a spectra query subsystem 116, or other components. Each of graph data generation subsystem 110, mass spectrum fragment determination subsystem 112, model training subsystem 114, and spectra query subsystem 116 may be configured to communicate with one another, one or more other devices, systems, and / or servers, using network 150 (e.g., the Internet, an Intranet). System 100 may also include one or more databases 140 (e.g., mass spectra database 142, training data database 144, model database 146, historical data database 148) used to store data for training machine learning models, storing machine learning models, or storing other data used by one orsf-5954611Docket No.: 226922001340 more components of system 100. This disclosure anticipates the use of one or more of each type of system and component thereof without necessarily deviating from the teachings of this disclosure.

[0063] Although not illustrated, other intermediary devices (e.g., data stores of a server connected to computing system 102) can also be used. The components of system 100 of FIG. 1 can be used in a variety of contexts where identifying and predicting mass spectra is performed. As an example, system 100 can be associated with a clinical environment where a user may input a mass spectrum as a query and may receive results indicating a top-k matching mass spectra. The user can review the mass spectra using client device 130. The user can provide additional information to computing system 102 that can be used to guide or direct the analysis.

[0064] In some embodiments, graph data generation subsystem 110 may be configured to generate and format graph data representing a chemical structure of a compound. In some embodiments, the graph data may represent the chemical structure of the compound. In one or more examples, graph data subsystem 110 may be configured to generate the graph data based on one or more simplified molecular-input line-entry system (SMILES) strings corresponding to the compound. An example SMILES string for the molecule Dinitrogen, having a chemical structure NŁN, is N#N. As another example, the SMILES string for the molecule Copper(II) Sulfate, having a chemical structure Cu2+SO42-, is [Cu+2].[O-]S(=O)(=O)[O-]. In one or more examples, graph data subsystem 110 may be configured to generate the graph data based on International Chemical Identifier (InChi) strings corresponding to the compound. As an example, the InChi string for Ethanol is InChI=1S / C2H6O / c1-2-3 / h3H,2H2,1H3.

[0065] In some embodiments, the graph data comprises information associated one or more experimental covariates. For example, the experimental covariates may comprise information of a type of mass spectrometer used to generate training data for training one or more machine learning models. As another example, the experimental covariates may comprise information of a setting of mass spectrometer 120 used to generate training data for training the one or more machine learning models. The setting of mass spectrometer 120 may include a collision energy, a collision gas, an ionization type, or other information, or combinations thereof. In one or more examples, the normalized collision energy may have a range of [0, 200].sf-5954611Docket No.: 226922001340

[0066] Spectra fragment determination subsystem 112 may be configured to generate an encoded representation of the chemical structure based on the graph data. In some embodiments, the encoded representation may comprise a plurality of graph features. As an example, the graph features may include nodes V, edges E, global attributes U, or other features. Nodes V may be connected to one another via edges E. Edges E may also be encoded to indicate whether they are directed or undirected. When representing molecules as graphs, the different atoms and bonds forming the molecule can be represented using a 3D object, where nodes V correspond to the atoms in the molecule and edges E correspond to the bonds (e.g., covalent bons) of the molecule. In some embodiments, an adjacency matrix indicating which atoms in a molecule are bonded to other atoms in the molecule may also be stored by the encoded representation of the chemical structure.

[0067] Spectra fragment determination subsystem 112 may further be configured to generate a plurality of mass spectrum fragments based on the encoded representation of the first chemical structure of the first compound and an identified set of compound fragments. The identified set of compound fragments may be stored, in mass spectra database 142, as a vocabulary. In one or more examples, the vocabulary may be associated with one or more machine learning models stored in model database 146. In some embodiments, the identified set of compound fragments may be identified based on a predetermined set of compound fragments. In one or more examples, the predetermined set of compound fragments comprise charged compound fragments and a derived set of neutral loss compound fragments. In one or more examples, the derived set of neutral loss compound fragments may be derived based on a mass of one or more neutral fragments of the predetermined set of compound fragments and a mass of the first chemical structure of the first compound. In one or more examples, first derived neutral loss compound fragment may be neutral.

[0068] In some embodiments, the mass spectrum fragments may be generated using one or more machine learning models. The machine learning models may be trained to generate a plurality of mass-to-charge values and a plurality of intensities. Each mass spectrum fragment may comprise one or more of the mass-to-charge values and one or more of the intensities. In some embodiments, spectra fragment determination subsystem 112 may be configured to generate, using the one or more machine learning models, a prediction of a mass spectrum of the compound based on the mass spectrum fragments.sf-5954611Docket No.: 226922001340

[0069] Model training subsystem 114 may be configured to train one or more machine learning models stored in model database 146 used to generate mass spectra fragments, predict mass spectra, transform a graphical representation of the chemical structure of the compound into one or more strings, or perform other functions. For instances, the strings may be ASCII strings which can be input to the machine learning models. In one or more examples, the chemical structure may be based on a SMILE string corresponding to the compound. In one or more examples, the chemical structure may be based on an InChi string corresponding to the compound. In some embodiments, model training subsystem 114 may be configured to execute one or more machine learning models that generate mass spectrum fragments from graph data representing a chemical structure of a compound. For example, graph data may represent the atoms and bonds of a molecule as nodes and edges of a graph. In this example, the machine learning models may include graph neural networks (GNNs) that are trained to receive the graph data as input, generate an encoded representation of the chemical structure based on the graph data, and generate mass spectrum fragments based on the encoded representation and an identified set of compound fragments.

[0070] Spectra query subsystem 116 may be configured to receive a query comprising a mass spectra and identify a top-k most similar mass spectra from mass spectra stored in mass spectra database 142. Mass spectra data 142 may store mass spectra derived experimentally as well as mass spectra generated using one or more processes described herein. In some embodiments, spectra query subsystem 116 may compute a similarity metric for some or all of the mass spectra stored in mass spectra database 142 as compared to the query spectrum. The similarity metric may indicate which derived / predicted mass spectra are most similar to the query spectrum. As an example, a weighted cosine similarity metric may be computed. In some embodiments, a top-k most similar spectra may be identified and provided to a user via client device 130. The value “k” may be configurable and may depend on (i) a number of mass spectra stored in mass spectra database 142, (ii) a complexity of the query spectrum, (iii) a threshold similarity (e.g., only mass spectra that, when compared to the query spectra, have a similarity that is greater than or equal to the threshold similarity), or (iv) user preference. In some examples, a user may adjust the value for “k” to view additional mass spectra or fewer mass spectra. Some example values for “k” include 3 or more, 10 or more, 50 or more, 100 or more, or other values.sf-5954611Docket No.: 226922001340

[0071] As described above, spectrum prediction refers to a process of predicting a mass spectrum of a compound. The predicted mass spectra can be added to a library of mass spectra, which may improve the ability of identifying new compounds by casting the spectrum prediction task as a search retrieval task. A query spectrum can be compared with mass spectra produced experimentally via a tandem mass spectrometer and generatively using one or more machine learning models. The query spectrum’s precursor compound can be identified based on which mass spectra in the library are determined to be most similar to the query spectrum.

[0072] In some embodiments, mass spectra may be generated using one or more machine learning models trained to predict a mass spectrum based on a molecular graph. An example molecular graph is depicted by precursor ion 300 of FIG. 3A. The mass spectrum comprises a variable-length set of peaks, each located at a (m / z) value. The particular (m / z) values relate to the molecular formula of the precursor ion that has been fragmented by mass spectrometer 120 of FIG. 1. The mass spectrum can be modeled as a probability distribution over molecular sub- formulas ښ(P), where P corresponds to the precursor. Modeling the mass spectrum as a probability distribution is more efficient than that of the bond-breaking process, which models a spectrum as a distribution over at worst exponentially many substructures of precursor P, ass many of the sub- formulas can be ruled out a priori. Furthermore, the techniques described herein preserve the precise (m / z) resolution of the bond breaking process, which is crucial to accurately predicting mass spectra.

[0073] One of the challenges of mass spectrum prediction for larger molecules is that the number of sub-formulas that are possible increases. To avoid having to enumerate every sub- formula, a select subset of the possible sub-formulas may be determined that explain a majority of small molecules. As an example, with reference to FIG. 4, plot 400 includes traces indicating a sum of all explained peaks’ heights within a given spectrum, averaged over all spectra, for a given vocabulary size. As illustrated by plot 400, almost all of the signal in small molecule mass spectra lies in peaks that can be explained by a relatively small number (~2%) of product ion and neutral loss formulas that frequently recur across spectra.

[0074] In some embodiments, molecular sub-formulas ښ(P) can be defined as:sf-5954611Docket No.: 226922001340

[0075] Here,^^represents a fixed set of frequent product ion formulas and ^ െ^^represents a variable set of ‘precursor-specific’ formulas determined by subtracting a fixed set of frequent neutral loss formulas^^from precursor ion ^. As a result of this approximation, the task of predicting a mass spectrum resolves to being a prediction of a probability of each of fixed setof frequent product ion formulas^^and fixed set of frequent neutral loss formulas^^.

[0076] The aforementioned technique leverages the fact that a product ion can be represented as its own formula or by a neutral loss formula relative to its precursor ion. Only including frequent product ion formulas produces poor results, particular as the mass of the compound increases.

[0077] In some embodiments, one or more machine learning models may be implemented to generate fixed set of frequent product ion formulas^^and fixed set of frequent neutral loss formulas^^. The model may be configured to list all product ion and neutral loss formulas yielded by mass decomposition of the training set, ranking them by the sum of the heights of all peaks to which each formula is assigned, and selecting the top-K highest rank among either type.

[0078] In some embodiments, the one or more machine learning models used to predict mass spectra fragments may include a graph neural network (GNN), however other graph machine learning models may be used in lieu of or in addition to the GNN. As an example, with reference to FIG. 5, graph neural network 504 may be trained to generate mass spectrum fragments 506 based on graph data 502. The training process may be performed using training data derived from experimental mass spectrum data. To tune one or more hyperparameters of graph neural network 504, a loss function may be minimized. For example, the loss-function used may be a peak-marginal cross entropy function:which is minimized with respect to the parameters of a neural network ^^^ή; ^^.

[0079] Thus, given graph data comprising a molecular graph G of a precursor ion formula P, graph neural network 504 may be trained to predict a probability ^^^for every formula f in the fixed vocabulary, which can be used to produce a spectrum S for the precursor ion. The per- formula probabilities may be summed within each observed peak across its compatible formulas to yield a predicted peak height, and the cross-entry loss may be calculated between the observedsf-5954611Docket No.: 226922001340 peaks (from training spectra) and the predicted peak heights across the entire spectrum may be minimized.

[0080] Graph neural network 504 may include an encoder 504a, a decoder 504b, and a projection layer 504c. Although a single encoder, decoder, projection layer are depicted, some examples include multiple encoders, decoders, and / or projection layers, or other layers. In some embodiments, graph neural network 504 may use a graph isomorphism network with edge and graph Laplacian features. Encoder 504a of graph neural network 504 may encode the graph data 502 into an encoded representation 508. In some embodiments, graph data 502 may comprise node features of the molecule, edge features, covariate features, and the top eigenvectors and eigenvalues of the graph Laplacian. In some examples, the 8 lowest-frequency eigenvalues may be used, which may be truncated or padded with zeros. In one or more examples, encoded representation 508 comprises embeddings. In some embodiments, a dense representation may be generated based on encoded representation 508 by attention pooling over the nodes and adding the embedded covariate features.

[0081] Each compound may be represented by its 2D graphical structure (e.g., as seen by precursor ion 300 of FIG.3A). In some examples, graph data 502 may represent the 2D graphical structure as G = (V, E), where V comprises the nodes of a graph representing atoms of the molecule and E comprises edges of the graph representing bonds between molecules. A virtual node may be added to include additional data associated with the graph. For example, graph data 502 may include node features V, edge features E, covariate features c, and the top eigenvectors and eigenvalues of the graph Laplacian vi and ^. In one or more examples, the node features and the edge features may be generated using an atom and bond feature generator. An example atom and bond feature generator is described in greater detail in “DGL-LifeSci: An Open-Source Toolkit for Deep Learning on Graphs in Life Science” to Li et al., 2021, the contents of which are hereby incorporated by reference in its entirety. In some embodiments, graph data 502 may include experimental parameters, such as normalized collision energy, precursor ion type, instrument model, and presence of isotopic peaks. Encoder 504a may be configured to generate encode this data to generate encoded representations 508 of the experimental parameters, such as the normalized collision energy, the precursor ion type, the instrument model, and the presence of isotopic peaks.sf-5954611Docket No.: 226922001340

[0082] In some embodiments, encoder 504a may be configured to embed nodes V, edges E, and covariate features c using a multilayer perceptron (MLP) block. In some embodiments, the top eigenvectors and eigenvalues vi and ^ may be transformed into node positional encodings using a sign invariant neural network. An example of a sign invariant neural network SignNet is described in “Sign and Basis Invariant Networks for Spectral Graph Representation Learning,” to Lim et al., 2022, the contents of which are incorporated herein by reference in their entireties. For example, SignNet may be used to encode the Laplacian features with ^ and ^ implemented as 2- layer stacked MLP blocks. In some embodiments, the embedded atom features and node positional encodings may be summed and, in along with the embedded bond features, may be passed to a stack of L message-passing layers to update the node representations and L MLP layers to update the edge representations. In one or more examples, encoder 504a may be an L = 6 layer encoder with 512 hyperparameters.

[0083] In some embodiments, the message-passing layer may use a graph isomorphism network with edge features. In one or more examples, two MLP blocks with a graphical normalization layer, such as, for example, GraphNorm, in place of layer normalization may be used as the network’s internal feed-forward network. Layer normalization may also be replaced with GraphNorm. GraphNorm normalizes the hidden representations across nodes in each individual graph with a learnable shift to avoid the expressiveness degradation while inheriting the acceleration effect of the shift operation. In some embodiments, node and edge updates may use residual connections, which improves training times.

[0084] In some embodiments, encoded representation 508 may be a dense representation of the molecule. In one or more examples, the dense representation may be generated by attention pooling over nodes, to which embedded covariate features may be embedded.

[0085] In some embodiments, decoder 504b of graph neural network 504 may be configured to generate a spectrum representation 510 based on encoded representation 508. In one or more examples, the spectrum representation comprises a histogram or (m / z, height) tuples each associated with the vocabulary. In some embodiments, decoder 504b may be implemented as a feed-forward network. In one or more examples, the feed-forward network comprises a stack of L* MLP blocks having residual connections. In one or more examples, decoder 504b may be an L* = 2 layer decoder including 1024 hyperparameters.sf-5954611Docket No.: 226922001340

[0086] In some embodiments, where encoder 504a and decoder 504b include, respectively, 512 and 1024 hyperparameters, graph neural network 504 may include 44.6 million trainable parameters.

[0087] In some embodiments, projection layer 504c may be configured to generate K-product ion probabilities 512 based on spectrum representation 510. The K-product ion probabilities 512 may include a set of logits zk for each of the K product ions or neutral loss formulas in the vocabulary.

[0088] Mass spectra fragments 506 may be generated based on K-product ion probabilities 512. For example, the probabilities may indicate which product ions or neutral loss formulas in the vocabulary are represented by graph data 502.

[0089] During training of graph neural network 504, mass spectra fragments 506 produced by graph neural network 504 may be compared to ground truth mass spectra fragments to calculate the peak-marginal cross entropy. For example, the training data used to train graph neural network 504 may include graph data associated with a molecule and an experimentally determined mass spectrum and / or mass spectra fragments. In some embodiments, the mass spectrum may be stored as tuples ofheight). In one or more examples, for a given molecule included in the training data, the peak-marginal cross entropy loss may be computed based on the experimentally determined mass spectrum and / or mass spectra fragments and the mass spectra fragments predicted by graph neural network 504. Hyperparameters of graph neural network 504 may be updated based on the computed peak-marginal cross entropy loss.

[0090] In some embodiments, additional corrections may be introduced to ensure that mass spectra fragments 506 are realistic. These corrections may relate to the possibility that some product ions may bind to ambient water or nitrogen molecules as adducts, which alters the mass of that fragment. The occurrences of this shifting can be included in the training data used to train graph neural network 504. To update the training data to account for the possible adducts, mass spectra fragments 506 may further include a logit for three different adduct states of each product ion formula in the vocabulary. The three states correspond to the original product ion f, the product ion bound with ambient water f + H20, and the product ion bound with nitrogen f + N2.

[0091] In some embodiments, additional corrections may be introduced to account for small peaks that can arise from higher isotopic states of the precursor ion P, at integral (m / z) shiftssf-5954611Docket No.: 226922001340 relative to the monoisotopic peak. Each predicted state, which now includes the original product ion f, the product ion bound with ambient water f + H20, and the product ion bound with nitrogen f + N2, may have a shared offset applied to each isotopic state.

[0092] In some embodiments, double counting may be performed as a result of the vocabulary including product ions and neutral loss formulas. In one or more example, a log 2 correction factor may be subtracted from the probability of each product ion and neutral loss to account for the possible double counting.

[0093] In some embodiments, graph neural network 504 may include an additional SoftMax layer that is used to generate mass spectra fragments ^^. After some or all of the corrections are applied, the SoftMax layer may be used to produce the heights for the various (m / z) peaks of mass spectra fragments 506.

[0094] In some embodiments, model training subsystem 114 may be configured to generate the training data based on the NIST-20 tandem spectral library. The NIST-20 tandem spectral library includes several spectra for a range of collision energies for each measured compound. Each spectrum comprises a list of (m / z, intensity, annotation) peak tuples, as well as metadata describing instrumental parameters and compound identity. In one or more examples, the annotation field includes a list of formula hypotheses per peak that were computed using mass decomposition. In some embodiments, the NIST-20 tandem spectral library may be restricted to HCD spectra with [M + H]+or [M - H]- precursor ions. In one or more examples, structures that are annotated as glycans or peptides or exceed 1000Da in mass (as these are not typically considered small molecules), or have atoms other than {C, H, N, O, P, S, F, Cl, Br, I} may be excluded.

[0095] In some embodiments, the training data may be split into training data, validation data, and test data. For example, an 85 / 5 / 10 structure-disjoint train / validation / test split may be generated by grouping spectra according to the connectivity substring of their ASCII chemical structure representation (e.g., InChi string) and assigning spectra to splits an entire group at a time.

[0096] As an example, with reference to FIG. 6, training process 600 may describe a process for training a graph neural network 604, in accordance with various embodiments. In some embodiments, training process 600 may generate or otherwise obtain training data 602 from training data database 144. In one or more examples, the training data may include graph datasf-5954611Docket No.: 226922001340 associated with a plurality of precursor ions, and corresponding experimentally derived mass spectrum associated with the precursor ions. The experimentally derived mass spectrum, for example, may include tuples of (m / z, intensity, annotation). In some embodiments, training data 602 may be input to graph neural network 604. In some embodiments, graph neural network 604 may correspond to a “to-be-trained” model retrieved from model database 146. Model database 146 may also store trained graph neural networks (e.g., graph neural network 604 subsequent to training process 600 being successfully executed).

[0097] Graph neutral network 604 may be configured to generate predicted mass spectrum fragments 606, using some or all of the processes described above with respect to FIG. 5. As an example, predicted mass spectra fragments 606 may include tuples of predicted (m / z, intensity, annotation). Predicted mass spectra fragments 606 and the predetermined mass spectra fragments included in training data 602 may be compared to compute loss 608. For example, the peak- marginal cross entry may be computed for loss 608. Based on loss 608, model training subsystem 114 may be configured to cause adjustments 610 to be made to some or all of the hyperparameters of graph neural network 604. Training process 600 may repeat for each molecule included in the training data and then may be validated and tested using the validation and test data, respectively.

[0098] As described above with respect to FIG.4, based on the training data used to train graph neural network 604 (e.g., NIST-20), most peaks relate to a small number of product ion and neutral loss formulas. In particular, plot 400 illustrates that a vocabulary of approximately 104product ion and neutral loss formulas can explain approximately 98% of ion counts detected in the training data. The present techniques further improve on existing techniques for mass spectra prediction by scaling better with molecular weight as compared to other mass spectra prediction techniques, such as, for example, bond-breaking. As an example, with respect to FIG.7, plots 700 depict a comparison of the mass spectra fragments predicted by graph neural network 504 trained using training process 600 for a select set of molecules. In plots 700, the blue lines represent the predicted spectra and the orange lines represent the ground truth spectra determined experimentally. The predicted spectra and ground-truth spectra are plotted against one another to illustrate the comparison of the intensity at the various (m / z) values.

[0099] In some embodiments, training process 600 may apply an all dropout at a rate of 0.1. Furthermore, a batch size of 512 and may be used along with the Adam optimizer with a learning rate of 5 × 10-4. In one or more examples, training process 600 may be performed forsf-5954611Docket No.: 226922001340 100 epochs. The model from the epoch with the lowest validation loss may be used as the trained machine learning model.

[0100] FIG. 8 shows an example flowchart of a method 800 for generating mass spectrum fragments, in accordance with various embodiments. Method 800 may be executed by one or more computing systems (e.g., one or more processors). For example, method 800 may be executed using one or more subsystems of computing system 102. In some embodiments, method 800 may begin at step 802.

[0101] At step 802, graph data representing a chemical structure of a compound may be obtained. In some embodiments, the graph data may include node features, edge features, covariate features, and top eigenvectors and eigenvalues of the graph. In some embodiments, the graph data may be generated. For example, the graph data representing the chemical structure of the compound may be generated based on one or more simplified molecular-input line-entry system (SMILES) strings corresponding to the compound. In some embodiments, the graph data representing the chemical structure of the compound may be generated based on one or more International Chemical Identifier (InChi) strings corresponding to the compound. In one or more examples, the graph data may comprise information associated one or more experimental covariates. The experimental covariates may comprise information of a type of mass spectrometer used to generate training data for training the machine learning models. The experimental covariates may comprise information of a setting of a mass spectrometer used to generate training data for training the machine learning models. In some examples, the setting of the mass spectrometer may include a collision energy, a collision gas, and / or an ionization type.

[0102] At step 804, an encoded representation of the chemical structure may be generated using one or more machine learning models. In some embodiments, the encoded representation may be generated using the machine learning models based on the graph data. In some embodiments, the encoded representation may comprise a plurality of graph features. In some embodiments, the encoded representation may be generated the machine learning model. In one or more examples, the machine learning model may comprise a graph neural network including an encoder. For example, encoder 504a of graph neural network 504 may encode graph data 502 into an encoded representation 508. In one or more examples, encoded representation 508 comprises embeddings. Some examples of the encoder may comprise a graph neural network (GNN), a graph convolutional network (GCN), or a graph isomorphism network (GIN), or othersf-5954611Docket No.: 226922001340 types of encoders. In some embodiments, a dense representation may be generated based on encoded representation 508 by attention pooling over the nodes and adding the embedded covariate features.

[0103] At step 806, a plurality of mass spectrum fragments may be generated using the machine learning models. In some embodiments, the mass spectrum fragments generated using the machine learning models may be based on the encoded representation of the chemical structure of the compound and an identified set of compound fragments. In some embodiments, the machine learning models may comprise a decoder. Some examples of the decoder may comprise a transformer-based decoder, a convolutional neural network (CNN) decoder, a feed-forward network decoder, a recurrent network decoder, or any combination thereof. The decoder can comprise a feed-forward motif, a recurrent connection, a self-attention layer, a convolutional layer, or any combination thereof. The decoder can be used to decode a spectrum representation from an encoded representation which can be readily interpreted by an experimenter. For example, the spectrum representation can comprise a series of mass-to-charge (i.e., m / z) values for one or more vocabulary molecules, e.g., molecules that can be predicted to derive from an input chemical structure, e.g., graph data, if a molecule of the input chemical structure was inputted into a mass spectrometer.

[0104] The decoder can decode the spectrum representation from the encoded representation, e.g., a low-dimensional representation, such as a matrix with preferred properties such as a matrix with preferred dimensions or elements. The decoder can comprise at least three to at least hundreds of layers in the network. At least one or more of the layers can comprise an activation layer, in which an activation function is applied to an input that is passed into the activation layer. The activation function can include a rectified linear unit (ReLU) function, a softmax function, a sigmoid function, or any other function that can normalize values, such as functions that normalize values between at least 0 and at most 1. At least one or more of the layers can comprise a pooling layer that comprises a pooling function, in which an input that is passed into the pooling layer is non-linearly downsampled, e.g., via an average pooling function or a max pooling function. The decoder can also comprise a skip connection across layers, e.g., a connection from an earlier layer to a non-adjacent layer. For the transformer-based decoder, the decoder can comprise a transformer, as described in Vaswani et al., (2017), “Attention is All You Need”, 31stConference on Neural Information Processing Systems. The transformer-based network can comprise ansf-5954611Docket No.: 226922001340 attention layer, such as a self-attention layer, and a pooling layer can comprise the attention layer. For the CNN decoder, at least one or more of the layers can comprise a convolutional layer, which can comprise a convolution function, in which an input into any one of the layers is transformed via a convolution filter, e.g., is convolved against a matrix, such as a kernel matrix. An activation layer can follow the convolutional layer. A pooling layer can follow the activation layer. For the feed-forward network decoder, the network can comprise a unidirectional flow of information from the original input to the output, e.g., the feed-forward network can lack recurrent connections. The feed-forward network can comprise skip connections across layers. For the recurrent network decoder, the network can comprise a recurrent connection, e.g., a connection between a later layer to an earlier layer, where the earlier layer may be adjacent or non-adjacent to the later layer.

[0105] In one or more examples, the identified set of compound fragments is stored as a vocabulary associated with the machine learning models. For example, the identified set of compound fragments may correspond to the K-product ions with which probabilities 512 is generated. In one or more examples, the identified set of compound fragments may be based on a predetermined set of compound fragments comprising charged compound fragments (e.g., fixed set of frequent product ion formulas^^) and a derived set of neutral loss compound fragments (e.g., fixed set of frequent neutral loss formulas^^from precursor ion^ ^. In some embodiments, method 800 may further include a step of selecting the predetermined set of compound fragments from a historical dataset of experimentally observed compound fragments associated with the compound. The historical dataset of experimentally observed compound fragments associated with a plurality of compounds may, in some embodiments, be stored in historical database 142. In some embodiments, selecting the predetermined set of compound fragments comprises determining a number of observations associated with each of the experimentally observed compound fragments from the historical dataset stored in historical database 148 and generating a ranking of the experimentally observed compound fragments based on the number of observations associated with each of the experimentally observed compound fragments. In some cases, the most- frequently observed compound fragments comprises a set of the most-probable compound fragments. In one or more examples, one or more of the experimentally observed compound fragments may be selected from the historical dataset stored in historical database 148 for the predetermined set of compound fragments, and thus are referred to as the set of most-probable compound fragments, based on the number of observations of each of the one or more of thesf-5954611Docket No.: 226922001340 experimentally observed compound fragments being greater than or equal to a threshold number of observations. In one or more examples, the predetermined set of compound fragments comprises 15,000 or less compound fragments, 10,000 or less compound fragments, 5,000 or less compound fragments, or 1,000 or less compound fragments. In one or more examples, the threshold number of observations may be 100 or more observations, 1,000 or more observations, 10,000 or more observations, 1,000,000 or more observations, or other quantities.

[0106] In some embodiments, method 800 may further include a step of generating the identified set of compound fragments. In one or more examples, generating the identified set of compound fragments may comprise selecting a set of compound fragments from a historical dataset of experimentally observed compound fragments associated with the compound, storing the selected set of compound fragments to the vocabulary, and updating the vocabulary based on the first derived set of neutral loss compound fragments, wherein the first identified set of compound fragments comprises the updated vocabulary.

[0107] The derived set of neutral loss compound fragments may be derived based on a mass of one or more neutral fragments of the predetermined set of compound fragments and a mass of the chemical structure of the compound, and each derived neutral loss compound fragment may be neutral. In some embodiments, method 800 may further include a step of generating the derived set of neutral loss compound fragments. Generating the derived set of neutral loss compound fragments may comprise subtracting a mass of one or more neutral loss compound fragments of the predetermined set of compound fragments from the mass of the chemical structure. In one or more examples, the one or more neutral loss compound fragments comprise a top-K set of neutral fragments from the mass of the chemical structure. In various embodiments, the phrase “top-K set" refers to a collection of the K elements with the highest rank or score based on a specific criterion, such as any of the criterion described above. In one or more examples, K equals 10,000 or less neutral fragments, 5,000 or less neutral fragments, or 1,000 or less neutral fragments.

[0108] In some embodiments, generating the mass spectrum fragments may comprise generating, using the machine learning models, a plurality of mass-to-charge values and a plurality of intensities. In one or more examples, each mass spectrum fragment comprises one or more of the plurality of mass-to-charge values and one or more of the plurality of intensities. In some embodiments, method 800 may further include a step of generating, using the machine learningsf-5954611Docket No.: 226922001340 models, a prediction of a mass spectrum of the compound based on the plurality of mass spectrum fragments. In some embodiments, method 800 may further include a step of identifying a set of most-probable compound fragments from the chemical structure. In one or more examples, the predetermined set of compound fragments comprises the set of most-probable compound fragments.

[0109] In some embodiments, the graph data, the chemical structure, and the compound described at step 802 may comprise first graph data, a first chemical structure, and a first compound, respectively. Method 800 may further include a step of obtaining second graph data representing a second chemical structure of a second compound and updating the vocabulary based on a second derived set of neutral loss compound fragments. In one or more examples, the second derived set of neutral loss compound fragments may be based on the mass of the one or more neutral fragments of the predetermined set of compound fragments and a mass of the second chemical structure. In one or more examples, the second derived set of neutral losses may be generated by subtracting the mass of the neutral fragments of the predetermined set of compound fragments from the mass of the second chemical structure. In some cases, the most-frequently observed compound fragments comprises a set of the most-probable compound fragments. In one or more examples, one or more of the experimentally observed compound fragments may be selected from the historical dataset stored in historical database 148 for the predetermined set of compound fragments, and thus are referred to as the set of most-probable compound fragments, based on the number of observations of each of the one or more of the experimentally observed compound fragments being greater than or equal to a threshold number of observations.

[0110] In some embodiments, method 800 may further comprise generating, using the machine learning models, a second encoded representation of the second chemical structure of the second compound based on the second graph data. In one or more examples, the second encoded representation of the second chemical structure comprises a second plurality of graph features. Method 800 may further comprise generating, using the machine learning models, a second plurality of mass spectrum fragments based on (i) the second encoded representation of the second chemical structure and an identified second set of compound fragments. In one or more examples, the identified second set of compound fragments may be based on the predetermined set of compound fragments and the second derived set of neutral loss compound fragments.sf-5954611Docket No.: 226922001340

[0111] In some embodiments, method 800 may further include a step of obtaining the compound. In some examples, the compound may be a natural product or a derivative thereof. In various embodiments, the phrase "natural product" means a chemical compound produced by a living organism, which may include plants, animals, fungi, and microorganisms. Non-limiting examples of suitable natural products are described in Harvey, A., Edrada-Ebel, R. & Quinn, R. The re-emergence of natural products for drug discovery in the genomics era. Nat Rev Drug Discov 14, 111–129 (2015). In some examples, the compound is of a plant extract or a derivative of the compound. In some examples, the compound may be a small molecule having a molecular weight of less than 2,000 Dalton (da).

[0112] In some embodiments, method 800 may be performed on a plurality of compounds to generate one or more mass spectra associated with each of the plurality of compound. The mass spectra of each compound may be compiled to generate the spectral library. In one or more examples, the compounds comprises a known compound. In one or more examples, the compounds comprises a theoretical compound.

[0113] In some embodiments, the one or more machine learning models can be trained by analyzing a mass spectrum using a spectral library, to predict a structure of a compound. The training of the machine learning model can comprise known methods in the art, including splitting a dataset into a training fraction and a validation fraction. The training can comprise cross- validating the machine learning model, such as by k-fold cross validation, iterative k-fold cross- validation, nested k-fold cross validation, or leave-p-out cross validation. The training of the machine learning algorithm, in the case that the machine learning algorithm is a neural network, e.g., a transformer-based network, a feed-forward network, a convolutional neural network, or a recurrent neural network, can comprise backpropagation. The backpropagation can include the adjusting of the weights of the neural network, based on the error-rate of the previous training iteration, while training the neural network.

[0114] FIG. 9 shows an example mass spectra query process 900, in accordance with various elements. In some embodiments, mass spectra query process 900 may be performed by spectra query subsystem 116. In some embodiments, mass spectra query process 900 may execute mass spectra query process 900 to identify a compound associated with a mass spectrum. In some examples, the mass spectrum may be analyzed using the spectral library to identify the compound associated with the mass spectrum. In some embodiments, a mass spectrometry technique maysf-5954611Docket No.: 226922001340 be performed to obtain the mass spectrum. As an example, spectra query subsystem 116 may receive a query spectrum 902 from mass spectrometer 120 and / or client device 130. Query spectrum 902 may be a 2D graph, a set of (m / z, intensity) tuples, or other data. Spectra query subsystem 116 may be configured to compute a similarity between query spectrum 902 and mass spectra 906 stored in mass spectra database 142. In one or more examples, mass spectra 906 may form part or all of the spectral library stored by mass spectra database 142. In some embodiments, the similarity may be computed by calculating a similarity metric, such as a weighted cosine metric:.

[0115] Based on the similarity metric computed, spectra query subsystem 116 may identify a top-K spectra 904 determined to be the most similar to query spectrum 902. Spectra query subsystem 116 may output top-K spectra 904 to client device 130. In some embodiments, spectra query subsystem 116 may generate an interface that allows for a user to adjust the value of K to view more or fewer of top-K spectra 904.

[0116] FIG.10 shows an example computing system 1000 used to implement one or more of the embodiments described herein. FIG. 10 illustrates an example computer system 1000. In particular embodiments, one or more computer systems 1000 perform one or more steps of one or more methods described or illustrated herein. In particular embodiments, one or more computer systems 1000 provide functionality described or illustrated herein. In particular embodiments, software running on one or more computer systems 1000 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 1000. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Moreover, reference to a computer system may encompass one or more computer systems, where appropriate.

[0117] This disclosure contemplates any suitable number of computer systems 1000. This disclosure contemplates computer system 1000 taking any suitable physical form. As example and not by way of limitation, computer system 1000 may be an embedded computer system, a system-sf-5954611Docket No.: 226922001340 on-chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on- module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these. Where appropriate, computer system 1000 may include one or more computer systems 1000; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 1000 may perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example, and not by way of limitation, one or more computer systems 1000 may perform in real time or in batch mode one or more steps of one or more methods described or illustrated herein. One or more computer systems 1000 may perform at different times or at different locations one or more steps of one or more methods described or illustrated herein, where appropriate.

[0118] In particular embodiments, computer system 1000 includes a processor 1002, memory 1004, storage 1006, an input / output (I / O) interface 1008, a communication interface 1010, and a bus 1012. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0119] In particular embodiments, processor 1002 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions, processor 1002 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 1004, or storage 1006; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 1004, or storage 1006. In particular embodiments, processor 1002 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 1002 including any suitable number of any suitable internal caches, where appropriate. As an example, and not by way of limitation, processor 1002 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches may be copies of instructions in memory 1004 or storage 1006, and the instruction cachessf-5954611Docket No.: 226922001340 may speed up retrieval of those instructions by processor 1002. Data in the data caches may be copies of data in memory 1004 or storage 1006 for instructions executing at processor 1002 to operate on; the results of previous instructions executed at processor 1002 for access by subsequent instructions executing at processor 1002 or for writing to memory 1004 or storage 1006; or other suitable data. The data caches may speed up read or write operations by processor 1002. The TLBs may speed up virtual-address translation for processor 1002. In particular embodiments, processor 1002 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 1002 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 1002 may include one or more arithmetic logic units (ALUs); be a multi-core processor; or include one or more processors 1002. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0120] In particular embodiments, memory 1004 includes main memory for storing instructions for processor 1002 to execute or data for processor 1002 to operate on. As an example, and not by way of limitation, computer system 1000 may load instructions from storage 1006 or another source (such as, for example, another computer system 1000) to memory 1004. Processor 1002 may then load the instructions from memory 1004 to an internal register or internal cache. To execute the instructions, processor 1002 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 1002 may write one or more results (which may be intermediate or final results) to the internal register or internal cache. Processor 1002 may then write one or more of those results to memory 1004. In particular embodiments, processor 1002 executes only instructions in one or more internal registers or internal caches or in memory 1004 (as opposed to storage 1006 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 1004 (as opposed to storage 1006 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple processor 1002 to memory 1004. Bus 1012 may include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) reside between processor 1002 and memory 1004 and facilitate accesses to memory 1004 requested by processor 1002. In particular embodiments, memory 1004 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Moreover,sf-5954611Docket No.: 226922001340 where appropriate, this RAM may be single-ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 1004 may include one or more memories 3404, where appropriate. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.

[0121] In particular embodiments, storage 1006 includes mass storage for data or instructions. As an example, and not by way of limitation, storage 1006 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Storage 1006 may include removable or non-removable (or fixed) media, where appropriate. Storage 1006 may be internal or external to computer system 1000, where appropriate. In particular embodiments, storage 1006 is non-volatile, solid-state memory. In particular embodiments, storage 1006 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 1006 taking any suitable physical form. Storage 1006 may include one or more storage control units facilitating communication between processor 1002 and storage 1006, where appropriate. Where appropriate, storage 1006 may include one or more storages 3406. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.

[0122] In particular embodiments, I / O interface 1008 includes hardware, software, or both, providing one or more interfaces for communication between computer system 1000 and one or more I / O devices. Computer system 1000 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and computer system 1000. As an example, and not by way of limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device or a combination of two or more of these. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interfaces 1008 for them. Where appropriate, I / O interface 1008 may include one or more device or software drivers enabling processor 1002 to drive one or more of these I / O devices. I / O interface 1008 may include one or more I / O interfacessf-5954611Docket No.: 226922001340 1008, where appropriate. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.

[0123] In particular embodiments, communication interface 1010 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between computer system 1000 and one or more other computer systems 1000 or one or more networks. As an example, and not by way of limitation, communication interface 1010 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 1010 for it. As an example, and not by way of limitation, computer system 1000 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 1000 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WI- MAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless network or a combination of two or more of these. Computer system 1000 may include any suitable communication interface 1010 for any of these networks, where appropriate. Communication interface 1010 may include one or more communication interfaces 1010, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0124] In particular embodiments, bus 1012 includes hardware, software, or both coupling components of computer system 1000 to each other. As an example and not by way of limitation, bus 1012 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low- pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or anothersf-5954611Docket No.: 226922001340 suitable bus or a combination of two or more of these. Bus 1012 may include one or more buses 1012, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.

[0125] Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field- programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid- state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.

[0126] Herein, “or” is inclusive and not exclusive, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A or B” means “A, B, or both,” unless expressly indicated otherwise or indicated otherwise by context. Moreover, “and” is both joint and several, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A and B” means “A and B, jointly or severally,” unless expressly indicated otherwise or indicated otherwise by context.

[0127] The scope of this disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the example embodiments described or illustrated herein that a person having ordinary skill in the art would comprehend. The scope of this disclosure is not limited to the example embodiments described or illustrated herein. Moreover, although this disclosure describes and illustrates respective embodiments herein as including particular components, elements, feature, functions, operations, or steps, any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that a person having ordinary skill in the art would comprehend. Furthermore, reference in the appended claims to an apparatus or system or a component of an apparatus or system being adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function encompasses that apparatus, system, component, whether or not it or that particular function is activated, turned on, or unlocked, as long as that apparatus, system, or component is so adapted, arranged, capable,sf-5954611Docket No.: 226922001340 configured, enabled, operable, or operative. Additionally, although this disclosure describes or illustrates particular embodiments as providing particular advantages, particular embodiments may provide none, some, or all of these advantages.

[0128] Certain aspects of the methods taught herein are discussed in a modular fashion. One of ordinary skill in the art will readily understand how the aspects of the present description can be combined to obtain any method encompassed by the teachings provided herein. The discussion of aspects of the methods in a modular fashion does not limit the scope of the description provided herein. A. Vocabularies

[0129] The vocabularies described herein may be based (such as assembled) on a diverse array of criteria. For example, in some embodiments, the vocabulary comprises about 100 compound fragments to about 50,000 compound fragments (such as any of about 100 compound fragments to about 25,000 compound fragments, about 100 compound fragments to about 15,000 compound fragments, about 5,000 compound fragments to about 15,000 compound fragments, or about 7,500 compound fragments to about 12,500 compound fragments) identified as the most common compound fragments (or comprising the most commonly observed chemical motifs) from one or more experimental data sets comprising known compounds and associated tandem mass spectra fragments. In some embodiments, the vocabulary comprises about 100,000 or fewer compound fragments (such as about any of 90,000 or fewer compound fragments, 80,00 or fewer compound fragments, 70,000 or fewer compound fragments, 60,000 or fewer compound fragments, 50,000 or fewer compound fragments, 40,000 or fewer compound fragments, 30,000 or fewer compound fragments, 20,000 or fewer compound fragments, 10,000 or fewer compound fragments, 9,000 or fewer compound fragments, 8,000 or fewer compound fragments, 7,000 or fewer compound fragments, 6,000 or fewer compound fragments, 5,000 or fewer compound fragments, 4,000 or fewer compound fragments, 3,000 or fewer compound fragments, 2,000 or fewer compound fragments, or 1,000 or fewer compound fragments) identified as the most common compound fragments (or comprising the most commonly observed chemical motifs) from one or more experimental data sets comprising known compounds and associated tandem mass spectra fragments. In some embodiments, the vocabulary comprises about any of 1,000 compoundsf-5954611Docket No.: 226922001340 fragments, 2,000 compound fragments, 3,000 compound fragments, 4,000 compound fragments, 5,000 compound fragments, 6,000 compound fragments, 7,000 compound fragments, 8,000 compound fragments, 9,000 compound fragments, 10,000 compound fragments, 11,000 compound fragments, 12,000 compound fragments, 13,000 compound fragments, 14,000 compound fragments, 15,000 compound fragments, 16,000 compound fragments, 17,000 compound fragments, 18,000 compound fragments, 19,000 compound fragments, 20,000 compound fragments, 25,000 compound fragments, 30,000 compound fragments, 35,000 compound fragments, 40,000 compound fragments, 45,000 compound fragments, or 50,000 compound fragments identified as the most common compound fragments (or comprising the most commonly observed chemical motifs) from one or more experimental data sets comprising known compounds and associated tandem mass spectra fragments.

[0130] In some embodiments, the vocabulary comprises context-specific compound fragments, such as fragments of compounds likely to be observed in a context, e.g., human metabolomics, plants and plant metabolomics, industrial processes, or petrochemicals.

[0131] In some embodiments, the vocabulary is customized with desired compound fragments. In some embodiments, the customization comprises adding desired compound fragments to an existing vocabulary established on a set criterion, such as a numerical cut-off and / or context-specific setting. B. Compounds

[0132] The methods provided herein are generally applicable to a vast array of compounds, including small molecular therapeutic compounds or candidate / precursors thereof. In some embodiments, the compound is a natural compound or a derivative thereof, such as an analog or metabolite. In some embodiments, the compound is of, such as obtained from, a plant (e.g., a plant extract), or a derivative, analog, or metabolite thereof, obtained from a plant. In some embodiments, the compound is from a plant material, or a processed material thereof, including from any aspect of a plant such as leaves, stems, fruit, flower, bark, wood, or root. In some embodiments, the compound is from a cultured plant cell, including an engineered plant cell. In some embodiments, the compound is a primary metabolite (which are directly required for plant growth), secondary (or specialized) metabolite (which mediate plant-environment interaction), orsf-5954611Docket No.: 226922001340 hormone (which regulate organismal processes and metabolism. E.g., Erb & Kliebenstein, Plant Physiol, 184, 2020, Atanasov et al., Nature Reviews Drug Discovery, 20, 2021, and Dias et al., Metabolites, 2, 2012, which are hereby incorporated herein by reference in their entirety. Techniques for obtaining compounds from a plant are known in the art, e.g., Nasim, Nucleus (Calcutta), 65, 2022, which is hereby incorporated herein by reference in its entirety.

[0133] In some embodiments, the compound comprises a small molecule having a molecular weight of less than 5,000 Daltons (da), such as less than any of 4,500 Daltons, 4,000 Daltons, 3,500 Daltons, 3,000 Daltons, 2,500 Daltons, 2,000 Daltons, 1,500 Daltons, 1,000 Daltons, 750 Daltons, 500 Daltons, or 250 Daltons.

[0134] In some embodiments, the compound is not a polypeptide, such as a peptide or a protein. In some embodiments, the compound is not a polypeptide comprising 3 or more amino acid residues. C. Mass spectrometry and covariates thereof

[0135] The methods provided herein, in certain aspects, use information of one or more experimental covariates associated with the analysis of a compound. In some embodiments, such experimental covariates are relevant to the vocabulary (i.e., experimental details associated with the experimental data used to generate each fragment of the vocabulary). In some embodiments, such experimental covariates are relevant to experimental conditions associated with (or in some embodiments associated with a planned mass spectrometry analysis) a predicted mass spectrum, e.g., fragmentation and / or m / z intensity is predicted under a certain experimental condition such as a set collision energy.

[0136] In some embodiments, the graph data corresponding to the chemical structure further comprises information of one or more experimental covariates. In some embodiments, the one or more experimental covariates comprises information of a type of mass spectrometer (such as discussed herein) used to generate training data for training one or more of the first machine learning model or the second machine learning model. In some embodiments, the one or more experimental covariates comprises information of a setting of a mass spectrometer (such as described herein) used to generate training data for training one or more of the first machine learning model or the second machine learning model. In some embodiments, the setting of thesf-5954611Docket No.: 226922001340 mass spectrometer includes one or more of a collision energy, a collision gas, or an ionization type. In some embodiments, the one or more experimental covariates comprises one or more of: a purification technique or an aspect thereof; an ionization technique or an aspect thereof (such as a setting involved with solvent flow), temperature (including temperature of the solvent and / or mass spectrometry interface), gas curtain, and / or voltage; a mass spectrometer or a component thereof, such as any front end additions including an ion funnel, any ion filtering / mobilization techniques, and / or detector types; or a fragmentation mechanism or an aspect thereof, such as collision energy, duration, and / or collision gas. In some embodiments, the collision type is collision-induced dissociation (CID), low-energy CID, sustained off-resonance irradiation collision-induced dissociation (SORI-CID), collision activated dissociation (CAD), higher-energy C-trap dissociation (HCD), electron capture dissociation (ECD), or electron-transfer dissociation (ETD).

[0137] In some embodiment, the method further comprises performing a mass spectrometry technique. In some embodiments, the mass spectrometry technique is performed to generate one or more experimental mass spectra. In some embodiments, the mass spectrometry technique is performed to validate a mass spectrum prediction of the methods described herein. In some embodiments, the method further comprises obtaining a compound and / or sample for the mass spectrometry technique. The present application contemplates a diverse array of mass spectrometry techniques for generating tandem mass spectra of a compound, including liquid chromatography mass spectrometry (LC-MS) techniques. In some embodiments, the mass spectrometry techniques provided herein provide (or are) one or more experimental covariate details used in the methods provided herein.

[0138] In some embodiment, the mass spectrometry technique comprises an ionization technique. Ionization techniques contemplated by the present application include techniques capable of charging polypeptides and peptide products. Thus, in some embodiments, the ionization technique is electrospray ionization. In some embodiments, the ionization technique is nano- electrospray ionization. In some embodiments, the ionization technique is atmospheric pressure chemical ionization. In some embodiments, the ionization technique is atmospheric pressure photoionization. In some embodiments, the ionization technique is matrix-assisted laser desorption ionization (MALDI). In some embodiment, the mass spectrometry technique comprises electrospray ionization, nano-electrospray ionization, or a matrix-assisted laser desorption ionization (MALDI) technique.sf-5954611Docket No.: 226922001340

[0139] Mass spectrometers contemplated by the present invention, to which an online liquid chromatography technique may be coupled, include high-resolution mass spectrometers and low- resolution mass spectrometers. Thus, in some embodiments, the mass spectrometer is a time-of- flight (TOF) mass spectrometer. In some embodiments, the mass spectrometer is a quadrupole time-of-flight (Q-TOF) mass spectrometer. In some embodiments, the mass spectrometer is a quadrupole ion trap time-of-flight (QIT-TOF) mass spectrometer. In some embodiments, the mass spectrometer is an ion trap. In some embodiments, the mass spectrometer is a single quadrupole. In some embodiments, the mass spectrometer is a triple quadrupole (QQQ). In some embodiments, the mass spectrometer is an orbitrap. In some embodiments, the mass spectrometer is a quadrupole orbitrap. In some embodiments, the mass spectrometer is a Fourier transform ion cyclotron resonance (FT) mass spectrometer. In some embodiments, the mass spectrometer is a quadrupole Fourier transform ion cyclotron resonance (Q-FT) mass spectrometer. In some embodiments, the mass spectrometry technique comprises positive ion mode. In some embodiments, the mass spectrometry technique comprises negative ion mode. In some embodiments, the mass spectrometry technique comprises a time-of-flight (TOF) mass spectrometry technique. In some embodiments, the mass spectrometry technique comprises a quadrupole time-of-flight (Q-TOF) mass spectrometry technique. In some embodiments, the mass spectrometry technique comprises an ion mobility mass spectrometry technique. In some embodiments a low-resolution mass spectrometry technique, such as an ion trap, or single or triple-quadrupole approach is appropriate.

[0140] In some embodiments, the LC-MS technique comprises separating a compound via a liquid chromatography technique. Liquid chromatography techniques contemplated by the present application include methods for separating compounds, such as small molecule compounds, and liquid chromatography techniques compatible with mass spectrometry techniques. In some embodiments, the liquid chromatography technique comprises a high-performance liquid chromatography technique. In some embodiments, the liquid chromatography technique comprises an ultra-high performance liquid chromatography technique. In some embodiments, the liquid chromatography technique comprises a high-flow liquid chromatography technique. In some embodiments, the liquid chromatography technique comprises a low-flow liquid chromatography technique, such as a micro-flow liquid chromatography technique or a nano-flow liquid chromatography technique. In some embodiments, the liquid chromatography technique comprises an online liquid chromatography technique coupled to a mass spectrometer. In some 42sf-5954611Docket No.: 226922001340 embodiments, the online liquid chromatography technique is a high-performance liquid chromatography technique. In some embodiments, the online liquid chromatography technique is an ultra-high performance liquid chromatography technique.

[0141] Other experimental steps useful for the methods provided herein are also contemplated. For example, in some embodiments, the method comprises obtaining a compound, such as an isolated or purified aliquot comprising the compound. D. Additional methods using a predicted spectrum or predicted spectra

[0142] In certain aspects, provided herein are methods leveraging a predicted spectrum or predicted spectra generating using the in silico prediction methods taught herein.

[0143] In some embodiments, provided is a method of generating a spectral library, the method comprising: performing a method of generating a mass spectrum fragment based on a chemical structure of a compound provided herein to generate one or more mass spectra associated with the compound, and compiling the one or more mass spectra associated with the compound to generate the spectral library.

[0144] In some embodiments, provided is a method of generating a spectral library, the method comprising: performing a method of generating a mass spectrum fragment based on a chemical structure of a compound provided herein for a plurality of compounds to generate one or more mass spectra associated with each compound, and compiling the one or more mass spectra associated with each compound to generate the spectral library.

[0145] In some embodiments, the compound or plurality of compounds comprise a known compound, such as a compound with a realized chemical structure known to exist in a composition. In some embodiments, the compound or plurality of compounds comprise a theoretical compound, such as a compound predicted as a possible precursor of or a generated from (such as a metabolite) a known compound.

[0146] In some embodiments, the spectral library is composed entirely of generated mass spectra, such as generated using the methods taught herein. In some embodiments, the spectral library is composed of generated mass spectra, such as generated using the methods taught herein, and experimentally obtained spectra. Spectral library may be of any size, such as from 1 spectrum to 1 million or more spectra.sf-5954611Docket No.: 226922001340

[0147] In some embodiments, provided herein is a method of identifying a compound associated with a mass spectrum, the method comprising: analyzing the mass spectrum using a spectral library described herein to identify the compound associated with the mass spectrum. In some embodiments, the method further comprises performing a mass spectrometry technique to obtain the mass spectrum associated with the compound. EXAMPLES

[0148] This example demonstrates a method of predicting tandem mass spectrum fragments based on a starting chemical structure using the methods taught herein.

[0149] With respect to FIG.7, plots 700 depict a comparison of the mass spectra fragments predicted by graph neural network 504 trained using training process 600 for a select set of molecules. In particular, plots 700 include a comparison of the predicted mass spectra fragments via graph neural network 504 to ground truth mass spectra fragments for progesterone (e.g., C21H30O2) and dydrogesterone (e.g., C21H28O2). To perform the prediction of mass spectrum fragments, graph data representing the chemical structure of progesterone and dydrogesterone was obtained. The graph data included nodes, edges, and other information associated with each compound. In some cases, the graph data included a SMILES string or an InChi string. The graph data representing the chemical structure of progesterone and dydrogesterone was input into a graph neural network. The graph neural network included an encoder-decoder architecture. The graph data was input to the encoder and the encoder generated an encoded representation of the chemical structures. The encoded representations were next input to the decoder and the decoder generated a plurality of mass spectrum fragments. For example, plots 700 of FIG.7 illustrate the predicted mass spectrum fragments generated from the graph data representing the 2D chemical structures of progesterone (e.g., C21H30O2) and dydrogesterone (e.g., C21H28O2). The decoder also used an identified set of compound fragments to generate the mass spectrum fragments. The set of compound fragments was stored as a vocabulary for the graph neural network. The identified set of compounds was formed based on a set known common fragments and a derived set of neutral loss compound fragments. The set of neutral loss compound fragments was derived by subtracting a mass of each neutral loss fragment from a mass of the chemical structure (e.g., one of progesterone (e.g., C21H30O2) and dydrogesterone (e.g., C21H28O2)).sf-5954611Docket No.: 226922001340

[0150] In plots 700, the blue lines represent the predicted spectra and the orange lines represent the ground truth spectra determined experimentally for progesterone, dydrogesterone, and diallyl phthalate. The predicted spectra and ground-truth spectra are plotted against one another to illustrate the comparison of the intensity at the various (m / z) values. As can be seen by plots 700, the graph neural network is able to predict the mass spectrum fragments for the chemical structure of progesterone and dydrogesterone. EXEMPLARY IMPLEMENTATIONS

[0151] Exemplary implementations of the methods and systems described herein include: 1. A method for generating mass spectrum fragments based on a chemical structure of a compound, the method comprising: obtaining first graph data representing a first chemical structure of a first compound; generating, using one or more machine learning models, a first encoded representation of the first chemical structure based the first graph data, wherein the first encoded representation comprises a first plurality of graph features; and generating, using the one or more machine learning models, a first plurality of mass spectrum fragments at least based on (i) the first encoded representation of the first chemical structure of the first compound and (ii) a first identified set of compound fragments, wherein: the first identified set of compound fragments is stored as a vocabulary associated with the one or more machine learning models, the first identified set of compound fragments is at least based on a predetermined set of compound fragments comprising first charged compound fragments and a first derived set of neutral loss compound fragments, the first derived set of neutral loss compound fragments is derived at least based on a mass of one or more neutral fragments of the predetermined set of compound fragments and a mass of the first chemical structure of the first compound, and each first derived neutral loss compound fragment is neutral. 2. The method of clause 1, wherein generating the first plurality of mass spectrum fragments comprises: generating, using the one or more machine learning models, a plurality of mass-to- charge values and a plurality of intensities, wherein each mass spectrum fragment comprises one or more of the plurality of mass-to-charge values and one or more of the plurality of intensities.sf-5954611Docket No.: 226922001340 3. The method of clause 2, further comprising: generating, using the one or more machine learning models, a prediction of a mass spectrum of the first compound at least based on the first plurality of mass spectrum fragments. 4. The method of any one of clauses 1-3, further comprising: identifying a set of most-probable compound fragments from the first chemical structure, wherein the predetermined set of compound fragments comprises the set of most-probable compound fragments. 5. The method of any one of clauses 1-4, further comprising: selecting the predetermined set of compound fragments from a historical dataset of experimentally observed compound fragments associated with the first compound. 6. The method of clause 5, wherein selecting comprises: determining a number of observations associated with each of the experimentally observed compound fragments from the historical dataset; generating a ranking of the experimentally observed compound fragments at least based on the number of observations associated with each of the experimentally observed compound fragments, wherein one or more of the experimentally observed compound fragments are selected from the historical dataset for the predetermined set of compound fragments at least based on the number of observations of each of the one or more of the experimentally observed compound fragments being greater than or equal to a threshold number of observations. 7. The method of any one of clauses 1-6, wherein the predetermined set of compound fragments comprises 15,000 or less compound fragments, 10,000 or less compound fragments, 5,000 or less compound fragments, or 1,000 or less compound fragments. 8. The method of any one of clauses 1-7, further comprising: generating the first identified set of compound fragments by: selecting a set of compound fragments from a historical dataset of experimentally observed compound fragments associated with the compound; storing the selected set of compound fragments to the vocabulary; and updating the vocabulary at least based on the first derived set of neutral loss compound fragments, wherein the first identified set of compound fragments comprises the updated vocabulary. 9. The method of any one of clauses 1-8, further comprising: generating the first derived set of neutral loss compound fragments by: subtracting a mass of one or more neutral loss compound fragments of the predetermined set of compound fragments from the mass of the first chemical structure.sf-5954611Docket No.: 226922001340 10. The method of clause 9, wherein the one or more neutral loss compound fragments comprise a top-K set of neutral fragments from the mass of the first chemical structure. 11. The method of clause 10, wherein K equals 10,000 or less neutral fragments, 5,000 or less neutral fragments, or 1,000 or less neutral fragments. 12. The method of any one of clauses 1-11, wherein the one or more machine learning models comprise an encoder. 13. The method of clause 12, wherein the encoder comprises at least one of: a graph neural network (GNN), a graph convolutional network (GCN), or a graph isomorphism network (GIN). 14. The method of any one of clauses 1-13, wherein the one or more machine learning models comprise a decoder. 15. The method of clause 14, wherein the decoder comprises at least one of a transformer-based decoder, a convolutional neural network (CNN) decoder, or a feed-forward network decoder. 16. The method of any one of clauses 1-15, further comprising: generating the first graph data representing the chemical structure at least based on one or more simplified molecular-input line- entry system (SMILES) strings corresponding to the compound. 17. The method of any one of clauses 1-16, further comprising: generating the first graph data representing the chemical structure at least based on one or more International Chemical Identifier (InChi) strings corresponding to the compound. 18. The method of any one of clauses 1-17, further comprising: obtaining second graph data representing a second chemical structure of a second compound different than the first compound; and updating the vocabulary based on a second derived set of neutral loss compound fragments, wherein the second derived set of neutral loss compound fragments is at least based on the mass of the one or more neutral fragments of the predetermined set of compound fragments and a mass of the second chemical structure. 19. The method of clause 18, wherein the second derived set of neutral loss compound fragments is generated by subtracting the mass of the one or more neutral fragments of the predetermined set of compound fragments from the mass of the second chemical structure. 20. The method of clause 19, further comprising: generating, using the one or more machine learning models, a second encoded representation of the second chemical structure of the secondsf-5954611Docket No.: 226922001340 compound at least based on the second graph data, wherein the second encoded representation of the second chemical structure comprises a second plurality of graph features; and generating, using the one or more machine learning models, a second plurality of mass spectrum fragments at least based on (i) the second encoded representation of the second chemical structure and an identified second set of compound fragments, wherein the identified second set of compound fragments is at least based on the predetermined set of compound fragments and the second derived set of neutral loss compound fragments. 21. The method of any one of clauses 1-20, further comprising: obtaining the compound. 22. The method of any one of clauses 1-21, wherein the compound is a natural product or a derivative thereof. 23. The method of any one of clauses 1-22, wherein the compound is of a plant extract or a derivative of the compound. 24. The method of any one of clauses 1-23, wherein the compound is a small molecule having a molecular weight of less than 2,000 Dalton (da). 25. The method of any one of clauses 1-24, wherein the first graph data comprises information associated with one or more experimental covariates. 26. The method of clause 25, wherein the one or more experimental covariates comprises information of a type of mass spectrometer used to generate training data for training the one or more machine learning models. 27. The method of clause 25, wherein the one or more experimental covariates comprises information of a setting of a mass spectrometer used to generate training data for training the one or more machine learning models. 28. The method of clause 27, wherein the setting of the mass spectrometer includes one or more of a collision energy, a collision gas, or an ionization type. 29. A method of generating a spectral library, the method comprising: performing the method of any one of clauses 1-28 on a plurality of compounds to generate one or more mass spectra associated with each of the plurality of compounds; and compiling the one or more mass spectra of each of the plurality of compounds to generate the spectral library. 30. The method of clause 29, wherein the plurality of compounds comprises a known compound.sf-5954611Docket No.: 226922001340 31. The method of clause 29, wherein the plurality of compounds comprises a theoretical compound. 32. A method of identifying a compound associated with a mass spectrum, the method comprising: analyzing the mass spectrum using the spectral library of any one of clauses 29-31 to identify the compound associated with the mass spectrum. 33. The method of clause 32, further comprising: performing a mass spectrometry technique to obtain the mass spectrum. 34. A method of identifying a compound associated with a mass spectrum, the method comprising: analyzing the mass spectrum using the spectral library of any one of claims 29-31 to train a machine learning model to predict a structure of the compound. 35. A system, comprising: one or more processors programmed to perform the method of any one of clauses 1-34. 36. A non-transitory computer-readable medium storing computer program instructions that, when executed by one or more processors of a computing system, effectuate operations comprising the method of any one of clauses 1-34.

[0152] It should be understood from the foregoing that, while particular implementations of the disclosed methods and systems have been illustrated and described, various modifications can be made thereto and are contemplated herein. It is also not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the preferable embodiments herein are not meant to be construed in a limiting sense. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. Various modifications in form and detail of the embodiments of the invention will be apparent to a person skilled in the art. It is therefore contemplated that the invention shall also cover any such modifications, variations and equivalents.sf-5954611

Claims

Docket No.: 226922001340 CLAIMS What is claimed is:

1. A method for generating mass spectrum fragments based on a chemical structure of a compound, the method comprising: obtaining first graph data representing a first chemical structure of a first compound; generating, using one or more machine learning models, a first encoded representation of the first chemical structure based the first graph data, wherein the first encoded representation comprises a first plurality of graph features; and generating, using the one or more machine learning models, a first plurality of mass spectrum fragments at least based on (i) the first encoded representation of the first chemical structure of the first compound and (ii) a first identified set of compound fragments, wherein: the first identified set of compound fragments is stored as a vocabulary associated with the one or more machine learning models, the first identified set of compound fragments is at least based on a predetermined set of compound fragments comprising first charged compound fragments and a first derived set of neutral loss compound fragments, the first derived set of neutral loss compound fragments is derived at least based on a mass of one or more neutral fragments of the predetermined set of compound fragments and a mass of the first chemical structure of the first compound, and each first derived neutral loss compound fragment is neutral.

2. The method of claim 1, wherein generating the first plurality of mass spectrum fragments comprises: generating, using the one or more machine learning models, a plurality of mass-to-charge values and a plurality of intensities, wherein each mass spectrum fragment comprises one or more of the plurality of mass-to-charge values and one or more of the plurality of intensities.

3. The method of claim 2, further comprising: generating, using the one or more machine learning models, a prediction of a mass spectrum of the first compound at least based on the first plurality of mass spectrum fragments.

4. The method of any one of claims 1-3, further comprising: identifying a set of most-probable compound fragments from the first chemical structure,sf-5954611Docket No.: 226922001340 wherein the predetermined set of compound fragments comprises the set of most-probable compound fragments.

5. The method of any one of claims 1-4, further comprising: selecting the predetermined set of compound fragments from a historical dataset of experimentally observed compound fragments associated with the first compound.

6. The method of claim 5, wherein selecting comprises: determining a number of observations associated with each of the experimentally observed compound fragments from the historical dataset; generating a ranking of the experimentally observed compound fragments at least based on the number of observations associated with each of the experimentally observed compound fragments, wherein one or more of the experimentally observed compound fragments are selected from the historical dataset for the predetermined set of compound fragments at least based on the number of observations of each of the one or more of the experimentally observed compound fragments being greater than or equal to a threshold number of observations.

7. The method of any one of claims 1-6, wherein the predetermined set of compound fragments comprises 15,000 or less compound fragments, 10,000 or less compound fragments, 5,000 or less compound fragments, or 1,000 or less compound fragments.

8. The method of any one of claims 1-7, further comprising: generating the first identified set of compound fragments by: selecting a set of compound fragments from a historical dataset of experimentally observed compound fragments associated with the compound; storing the selected set of compound fragments to the vocabulary; and updating the vocabulary at least based on the first derived set of neutral loss compound fragments, wherein the first identified set of compound fragments comprises the updated vocabulary.

9. The method of any one of claims 1-8, further comprising: generating the first derived set of neutral loss compound fragments by: subtracting a mass of one or more neutral loss compound fragments of the predetermined set of compound fragments from the mass of the first chemical structure.sf-5954611Docket No.: 226922001340 10. The method of claim 9, wherein the one or more neutral loss compound fragments comprise a top-K set of neutral fragments from the mass of the first chemical structure.

11. The method of claim 10, wherein K equals 10,000 or less neutral fragments, 5,000 or less neutral fragments, or 1,000 or less neutral fragments.

12. The method of any one of claims 1-11, wherein the one or more machine learning models comprise an encoder.

13. The method of claim 12, wherein the encoder comprises at least one of: a graph neural network (GNN), a graph convolutional network (GCN), or a graph isomorphism network (GIN).

14. The method of any one of claims 1-13, wherein the one or more machine learning models comprise a decoder.

15. The method of claim 14, wherein the decoder comprises at least one of a transformer- based decoder, a convolutional neural network (CNN) decoder, or a feed-forward network decoder.

16. The method of any one of claims 1-15, further comprising: generating the first graph data representing the chemical structure at least based on one or more simplified molecular-input line-entry system (SMILES) strings corresponding to the compound.

17. The method of any one of claims 1-16, further comprising: generating the first graph data representing the chemical structure at least based on one or more International Chemical Identifier (InChi) strings corresponding to the compound.

18. The method of any one of claims 1-17, further comprising: obtaining second graph data representing a second chemical structure of a second compound different than the first compound; and updating the vocabulary based on a second derived set of neutral loss compound fragments, wherein the second derived set of neutral loss compound fragments is at least based on the mass of the one or more neutral fragments of the predetermined set of compound fragments and a mass of the second chemical structure.sf-5954611Docket No.: 226922001340 19. The method of claim 18, wherein the second derived set of neutral loss compound fragments is generated by subtracting the mass of the one or more neutral fragments of the predetermined set of compound fragments from the mass of the second chemical structure.

20. The method of claim 19, further comprising: generating, using the one or more machine learning models, a second encoded representation of the second chemical structure of the second compound at least based on the second graph data, wherein the second encoded representation of the second chemical structure comprises a second plurality of graph features; and generating, using the one or more machine learning models, a second plurality of mass spectrum fragments at least based on (i) the second encoded representation of the second chemical structure and an identified second set of compound fragments, wherein the identified second set of compound fragments is at least based on the predetermined set of compound fragments and the second derived set of neutral loss compound fragments.

21. The method of any one of claims 1-20, further comprising: obtaining the compound.

22. The method of any one of claims 1-21, wherein the compound is a natural product or a derivative thereof.

23. The method of any one of claims 1-22, wherein the compound is of a plant extract or a derivative of the compound.

24. The method of any one of claims 1-23, wherein the compound is a small molecule having a molecular weight of less than 2,000 Dalton (da).

25. The method of any one of claims 1-24, wherein the first graph data comprises information associated with one or more experimental covariates.

26. The method of claim 25, wherein the one or more experimental covariates comprises information of a type of mass spectrometer used to generate training data for training the one or more machine learning models.sf-5954611Docket No.: 226922001340 27. The method of claim 25, wherein the one or more experimental covariates comprises information of a setting of a mass spectrometer used to generate training data for training the one or more machine learning models.

28. The method of claim 27, wherein the setting of the mass spectrometer includes one or more of a collision energy, a collision gas, or an ionization type.

29. A method of generating a spectral library, the method comprising: performing the method of any one of claims 1-28 on a plurality of compounds to generate one or more mass spectra associated with each of the plurality of compounds; and compiling the one or more mass spectra of each of the plurality of compounds to generate the spectral library.

30. The method of claim 29, wherein the plurality of compounds comprises a known compound.

31. The method of claim 29, wherein the plurality of compounds comprises a theoretical compound.

32. A method of identifying a compound associated with a mass spectrum, the method comprising: analyzing the mass spectrum using the spectral library of any one of claims 29-31 to identify the compound associated with the mass spectrum.

33. The method of claim 32, further comprising: performing a mass spectrometry technique to obtain the mass spectrum.

34. A method of identifying a compound associated with a mass spectrum, the method comprising: analyzing the mass spectrum using the spectral library generated by any one of claims 29-31 to train a machine learning model to predict a structure of the compound.

35. A system, comprising: one or more processors programmed to perform the method of any one of claims 1-34.sf-5954611Docket No.: 226922001340 36. A non-transitory computer-readable medium storing computer program instructions that, when executed by one or more processors of a computing system, effectuate operations comprising the method of any one of claims 1-34.sf-5954611