Method for analysis of a chromatogram of samples containing isomers
The method accurately assigns isomers in chromatograms by predicting unique fragments and using mass spectral data, enhancing the analysis of isomer-containing samples.
Patent Information
- Application Number
- PCT/EP2024/067349
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-12-26
AI Technical Summary
Existing methods struggle to accurately assign isomers, which have the same mass but different structures, to their corresponding peaks in a chromatogram.
A method that involves obtaining a chromatogram and associated mass spectral data, predicting fragments of candidate chemical structures, and assigning isomers to peaks based on unique m/z ratios of these fragments, using a computer program and apparatus comprising a chromatography instrument and mass spectrometer.
Enables accurate assignment of isomers to their respective chromatogram peaks, improving the analysis of samples containing isomers.
Smart Images

Figure EP2024067349_26122025_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR ANALYSIS OF A CHROMATOGRAM OF SAMPLES CONTAINING ISOMERS
[0002] Field
[0003] The claimed invention relates to a method for analysis of a chromatogram. Particularly, a method that enables assigning of isomers to peaks of a chromatogram.
[0004] Background
[0005] In known methods of assigning corresponding compounds to peaks of a chromatogram, one or more mass spectra associated with each peak would be used to ascertain which compound to assign to which peak. However, such known techniques do not readily enable assignment of isomers (compounds having the same mass but different structures) to their corresponding peaks.
[0006] Samples obtained by derivatization of a compound often contain isomers. In chromatography, derivatization of a compound is often performed before input to the chromatography column to improve, for example, the chromatographic separation and / or thermal stability of the compound. As is well known in the art, derivatization is the process of chemically altering a compound by reacting the compound with a derivatization agent. Derivatization can involve, for example, silylation, alkylation or acylation reactions. During derivatization, the derivatization agent reacts with a functional group in the compound, for example active hydrogens in the case of silylation. Such functional groups may be referred to as active functional groups. If the compound contains multiple functional groups, then derivatization of the compound can result in a mixture of derivatives where the derivatized functional group may be in a different position in each derivative. Such derivatives may also be isomers having the same mass but different structural formula. As discussed above, it is difficult to assign isomers to their respective chromatogram peaks.
[0007] Summary
[0008] Against this background, there is provided a method of analysing a chromatogram of a sample, the method comprising: obtaining a chromatogram of the sample, the sample comprising a mixture of a plurality of compounds, wherein at least two of the compounds are isomers, the chromatogram comprising a plurality chromatogram peaks; obtaining associated mass spectral data for each chromatogram peak of the plurality of chromatogram peaks; predicting fragments of each candidate chemical structure of a first plurality of candidate chemical structures and identifying a m / z ratio of at least one unique fragment of each candidate chemical structure of the first plurality of candidate chemical structures, wherein at least two of the candidate chemical structures of the first plurality of candidate chemical structures are candidate isomers; and for one or more of the candidate chemical structure(s) of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak of the plurality of chromatogram peaks based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio(s) of the at least one unique fragment(s) of the candidate chemical structure, wherein at least one of the assigned candidate chemical structure(s) is one of the candidate isomers.
[0009] A computer program comprising instructions, that when executed by a processor, cause the processor to perform any method herein disclosed is also provided. An apparatus comprising: a chromatography instrument configured to obtain the chromatogram; a mass spectrometer configured to obtain the associated mass spectral data for each chromatogram peak and a data analysis system configured to carry out any method herein disclosed using the chromatogram obtained by the chromatography instrument and mass spectral data obtained by the mass spectrometer is also provided.
[0010] The method of the claimed invention achieves improved analysis of a chromatogram of a sample comprising isomers.
[0011] The claimed invention is particularly advantageous for enabling the correct isomers to be assigned to the corresponding peaks of a chromatogram of a sample comprising a mixture of isomers. The method may be applicable to chromatograms obtained using for example, gas chromatography, liquid chromatography, ion chromatography.
[0012] In known methods of assignment of compounds to peaks of a chromatogram, mass spectra associated with each peak would be used to ascertain which compound to assign to which peak. However, such known techniques do not readily facilitate the identification of isomers that have the same mass but different structures. The method of the claimed invention resolves the ambiguity as to which isomer should be assigned to which chromatogram peak.
[0013] Herein, a chromatogram peak refers to a peak in a chromatogram where a chromatogram is a plot of detector response (detector signal intensity) against retention time. A chromatogram peak typically has a Gaussian shaped profile, or can be assumed to have a Gaussian shaped profile.
[0014] Herein, a m / z peak (also known as a mass peak) refers to a peak in a mass spectrum. Candidate chemical structures are compounds that are predicted to be compounds present in the sample and so correspond to peaks of the chromatogram. Candidate isomers are candidate chemical structures that are also isomers i.e. having the same molecular formula but a different structural formula. The candidate chemical structures may be defined according to their structural formulas. The candidate chemical structures may all be isomers (referred to herein as candidate isomers).
[0015] A unique fragment of a candidate chemical structure of the first plurality of candidate chemical structures refers to a fragment predicted to be generated on fragmentation of a candidate chemical structure that is not predicted to be generated on fragmentation of another candidate chemical structure within the first plurality of candidate chemical structures. At least two of the candidate chemical structure(s) within the first plurality of candidate chemical structures are candidate isomers. Candidate isomers are candidate chemical structures having the same molecular formula but different structural formula. A unique fragment of a candidate isomer would be a fragment predicted to be generated on fragmentation of the candidate isomer that would not be predicted on fragmentation of the other candidate chemical structure(s) within the first plurality of candidate chemical structures (including the other candidate isomer(s) within the first plurality of candidate chemical structures). A unique fragment of a candidate chemical structure has a m / z ratio that is different from the fragments predicted for the other candidate chemical structure(s). Optionally, a plurality of unique fragments are predicted for each candidate chemical structure.
[0016] Associated mass spectral data for a chromatogram peak refers to the mass spectral data corresponding to the chromatogram peak. The mass spectral data associated with a chromatogram peak is the mass spectral data of the compounds eluted from the chromatography column at a respective value of the retention time corresponding to that chromatogram peak.
[0017] The associated mass spectral data for each peak may comprise one or more MSnmass spectra where n > 1. In other words, the mass spectral data may comprise mass spectra generated in respect of the unfragmented ionised compounds (MS1spectra) and / or generated in respect of fragments of the ionised compounds (e.g. MSnspectra where n>2). The associated mass spectral data may be one of a plurality of mass spectra associated with the chromatogram peak. For example, each peak may have an associated MS1spectrum and MS2spectrum. Each peak may have a plurality of multi-stage MSnassociated mass spectra that may form a mass spectral tree. The mass spectral data may be a list of identified m / z peaks in the one or more MSnmass spectra where n > 1 .
[0018] According to the method, a candidate chemical structure may only be assigned to one chromatogram peak (the most probable chromatogram peak) such that a chromatogram peak only has a single candidate chemical structure assigned thereto. The most probable chromatogram peak is the chromatogram peak that is most likely to correspond to the candidate chemical structure.
[0019] The step of assigning a candidate chemical structure to a chromatogram peak may be performed for each of the candidate chemical structures so that each candidate chemical structure is assigned to a chromatogram peak (the most probable chromatogram peak for that candidate chemical structure).
[0020] The first plurality of candidate chemical structures includes at least two candidate isomers. The step of assigning a candidate chemical structure to a chromatogram peak of the plurality of chromatogram peaks may be performed for each of the candidate isomers so that each candidate isomer is assigned to a chromatogram peak that is the most probable chromatogram peak for that candidate isomer. This is such that each candidate isomer would be assigned to a single chromatogram peak.
[0021] Optionally, the step of assigning the candidate chemical structure to a chromatogram peak of the plurality of chromatogram peaks, comprises determining a score for each chromatogram peak indicating the extent to which the candidate chemical structure correlates to the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak having the score indicating the greatest correlation, wherein the score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data that correspond to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure. The step of assigning the candidate chemical structure to a most probable chromatogram peak is performed in respect of at least one of the candidate isomers such that the candidate isomer is assigned to the most probable chromatogram peak.
[0022] The score may be calculated based on identifying, in the associated mass spectral data, a number and / or intensity of the one or more m / z peak(s) corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure.
[0023] The score may be calculated based on a ratio of the sum of intensities of m / z peak(s) corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure to the sum of intensities of all m / z peaks in the associated mass spectral data.
[0024] The score may be calculated based on a ratio of a number of m / z peak(s) corresponding to the m / z ratio(s) of unique fragment(s) of the candidate chemical structure to a number of all m / z peaks in the associated mass spectral data. Optionally, the method may comprise obtaining a list of candidate chemical structures (for the plurality of compounds), wherein the list of candidate chemical structures includes the first plurality of candidate chemical structures. The list of candidate chemical structures may include candidate chemical structures in addition to the first plurality of candidate chemical structures for which fragments are then predicted. The list of candidate chemical structures may include the first plurality of candidate chemical structures and a second plurality of candidate chemical structures. In such an arrangement, the method may comprise obtaining a list of candidate chemical structures and predicting fragments (and identifying unique fragment(s)) of a subset of that list of candidate chemical structures (the subset referred to as the first plurality of candidate chemical structures herein). Alternatively, the list of candidate chemical structures may consist of (comprise only) the first plurality of candidate chemical structures.
[0025] Optionally, the mixture of the plurality of compounds consists (comprises only) of isomers. Similarly, the first plurality of candidate chemical structures may consist of (comprise only) candidate isomers.
[0026] Optionally, the list of candidate chemical structures are obtained using a database, optionally wherein the database is a mass spectral library. Optionally the list of candidate chemical structures are obtained based on (i) the associated mass spectral data of each chromatogram peak; and / or based on (ii) knowledge of a precursor to one or more compound(s) of the plurality of compounds or knowledge of a molecular formula of one or more compound(s) of the plurality of compounds.
[0027] In one example, some candidate chemical structures within the list of candidate chemical structures may be obtained based on (i) the associated mass spectral data of each chromatogram peak and / or some candidate chemical structures may be obtained based on (ii) knowledge of a precursor to one or more compound(s) of the plurality of compounds and / or some of the candidate chemical structures may be obtained based on knowledge of a molecular formula of one or more compound(s) of the plurality of compounds.
[0028] Optionally the step of obtaining the list of candidate chemical structures based on the associated mass spectral data of each chromatogram peak comprises searching the database for candidate chemical structures having mass spectral data corresponding to the mass spectral data associated with each chromatogram peak. Optionally, the associated mass spectral data for each chromatogram peak comprises an associated mass spectrum, wherein searching the database for candidate chemical structures having mass spectral data corresponding to the mass spectral data associated with each chromatogram peak comprises: identifying one or more m / z peak(s) in the associated mass spectrum; and identifying a compound in the database as a candidate chemical structure when one or more m / z ratio(s) associated with the compound, for example m / z peak(s) of ion(s) of the compound, in the database corresponds with one or more m / z ratio(s) of the identified m / z peak(s) in the associated mass spectrum. The tolerance range may be defined by the accuracy of the mass spectral data.
[0029] In accordance with an embodiment of the method where the list of candidate chemical structures are obtained based the associated mass spectral data of each chromatogram peak, the method may further comprise, before at least the step of assigning the candidate chemical structure to a chromatogram peak (and optionally also before the step of predicting fragments of each candidate chemical structure of a first plurality of candidate chemical structures and identifying a m / z ratio of at least one unique fragment of each candidate chemical structure of the first plurality of candidate chemical structures), the step of: allocating candidate chemical structures from the list of candidate chemical structures (the list of candidate chemical structures including the first plurality of candidate chemical structures) to one or more of the chromatogram peak(s) based on the associated mass spectral data of the chromatogram peaks. In other words the method may comprise allocating multiple, optionally all, candidate chemical structures from the list of candidate chemical structures to chromatogram peaks where those allocated candidate chemical structures include the first plurality of candidate chemical structures. According to this method, each candidate isomer would be allocated to multiple chromatogram peaks, since the allocation is performed based on the associated mass spectral data of the chromatogram peaks.
[0030] In accordance with this embodiment of the method, candidate chemical structures (which may be candidate isomers) may be allocated to multiple chromatogram peaks and a chromatogram peak may have multiple candidate chemical structures (which may be candidate isomers) allocated thereto. The term “allocating” used herein refers to associating the candidate chemical structures with possible chromatogram peaks. In contrast, the term “assigning” used herein refers to associating the candidate chemical structures (which may be candidate isomers) with their most probable chromatogram peak. A single candidate chemical structure (which may be a candidate isomer) is assigned a single chromatogram peak. Typically, candidate isomers will be allocated to multiple chromatogram peaks based on the associated mass spectral data of the chromatogram peaks. However, each candidate isomer will only be assigned to a single chromatogram peak, which is the most probable chromatogram peak. Optionally, the first plurality of candidate chemical structures consists of isomers having a first molecular formula.
[0031] In an embodiment where the first plurality of candidate chemical structures consists of isomers having a first molecular formula, the method may further comprise grouping chromatogram peaks of the plurality of chromatogram peaks having one or more candidate chemical structure(s) of the first plurality of candidate chemical structures allocated thereto into a first group. In this embodiment, the step of, for one or more candidate chemical structure(s) of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak, comprises assigning the candidate chemical structure to a chromatogram peak within the first group based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratios corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure.
[0032] Optionally the step of assigning the candidate chemical structure of the first plurality to a (most probable) chromatogram peak within the group comprises determining a score for each chromatogram peak within the group indicating the extent to which the candidate chemical structure within the group correlates to the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak within the first group having the score indicating the greatest correlation, wherein the score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data that correspond to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure. The scoring methods described herein equally apply to this embodiment.
[0033] For an embodiment where the list of candidate chemical structures further comprises a second plurality of candidate chemical structures, wherein each candidate chemical structure of the second plurality of candidate chemical structures has a second molecular formula (i.e. the second plurality of candidate chemical structures consists of candidate chemical structures that are either the same or are isomers (candidate isomers)), the second molecular formula being different from the first molecular formula, the method may further comprise grouping chromatogram peaks having one or more candidate chemical structure(s) of the second plurality candidate chemical structures allocated thereto into a second group. The method further comprises predicting fragments of each candidate chemical structure of the second plurality of candidate chemical structures and identifying a m / z ratio of at least one unique fragment of each candidate chemical structure of the second plurality of candidate chemical structures, wherein at least two of the candidate chemical structures of the second plurality of candidate chemical structures are candidate isomers; and for one or more of the candidate chemical structure(s) of the second plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak of the plurality of chromatogram peaks within the second group based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio(s) of the at least one unique fragment(s) of the candidate chemical structure, wherein at least one of the assigned candidate chemical structure(s) is one of the candidate isomers.
[0034] Optionally, in an embodiment where the step of assigning the candidate chemical structure of the first plurality to a (most probable) chromatogram peak within the group comprises determining a score for each chromatogram peak within the group indicating the extent to which the candidate chemical structure within the group correlates to the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak within the first group having the score indicating the greatest correlation, wherein the score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data that correspond to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure. The scoring methods described herein equally apply to this embodiment.
[0035] In other words, the method may further comprise, after the step of allocating candidate chemical structures from the list of candidate chemical structures to one or more of the chromatogram peak(s), grouping chromatogram peaks into one or more groups such that each group contains chromatogram peaks with allocated candidate chemical structures having the same molecular formula.
[0036] For clarity, an embodiment where the list of candidate chemical structures includes a first plurality of candidate chemical structures having a first molecular formula and a second plurality of candidate chemical structures having a second molecular formula can be set out as: a method of analysing a chromatogram of a sample, the method comprising: obtaining a chromatogram of the sample, the sample comprising a mixture of a plurality of compounds, wherein at least two of the compounds are isomers, the chromatogram comprising a plurality of chromatogram peaks; obtaining associated mass spectral data for each chromatogram peak; obtaining a list of candidate chemical structures based on the associated mass spectral data of each chromatogram peak, where the list of candidate chemical structures comprise a first plurality of candidate chemical structures having a first molecular formula, wherein at least two of the candidate chemical structures are candidate isomers; allocating candidate chemical structures from the list of candidate chemical structures, including the candidate isomers, to one or more of the chromatogram peaks based on the associated mass spectral data of the chromatogram peak; grouping together chromatogram peaks of the plurality of chromatogram peaks having one or more candidate chemical structure(s) of the first plurality of candidate chemical structures allocated thereto into a first group; predicting fragments of each candidate chemical structure of the first plurality of candidate chemical structures; and identifying a m / z ratio of at least one unique fragment of each candidate chemical structure of the first plurality of candidate chemical structures; and for one or more of the candidate chemical structure(s) of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak within the group based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio(s) of the at least one unique fragment(s) of the candidate chemical structure, wherein least one of the assigned candidate chemical structures is a candidate isomer.
[0037] In accordance with an embodiment of the method, the list of candidate chemical structures are obtained based on knowledge of a precursor to one or more compound(s) of the plurality of compounds. In one exemplary embodiment, the first plurality of candidate chemical structures may be predicted derivatives of the known precursor. At least two of the predicted derivatives are candidate isomers. Optionally, the predicted derivatives are determined by generating derivatives of the known precursor in-silico (by simulation or computer modelling).
[0038] For clarity, such an embodiment can be set out as a method of analysing a chromatogram of a sample, the method comprising: obtaining a chromatogram of the sample, the sample comprising a mixture of a plurality of compounds, wherein the plurality of compounds comprise derivatives of a known precursor, wherein at least two of the derivatives are isomers, the chromatogram comprising a plurality of chromatogram peaks; obtaining associated mass spectral data for each chromatogram peak; obtaining a list of candidate chemical structures, wherein the list of candidate chemical structures comprise a first plurality of candidate chemical structures that are predicted derivatives of the known precursor, wherein at least two of the predicted derivatives are candidate isomers; predicting fragments of the first plurality of candidate chemical structures and identifying a m / z ratio of at least one unique fragment for each candidate chemical structure of the first plurality of candidate chemical structures; and for one or more of the candidate chemical structure(s) of the first plurality of candidate chemical structures, assigning the candidate isomer to a most probable chromatogram peak of the plurality of chromatogram peaks based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio of the at least one unique fragment(s) of the candidate chemical structure.
[0039] Optionally, all of the predicted derivatives are candidate isomers.
[0040] Optionally the first plurality of candidate chemical structures consists of the predicted derivatives. Optionally, the list of candidate chemical structures consists of the predicted derivatives and, optionally, the known precursor.
[0041] Optionally, the method comprises performing derivatisation of the known precursor to obtain the sample and performing chromatography of the sample to obtain the chromatogram.
[0042] Brief Description of Figures
[0043] Figure 1 schematically illustrates an example apparatus for carrying out the methods described herein.
[0044] Figure 2 shows a flow diagram indicating steps of a method in accordance with an embodiment of the disclosure.
[0045] Figure 3 is an exemplary embodiment of the method of Figure 1 where the sample used to obtain the chromatogram comprises a mixture of compounds where candidate structural formulas for the compounds are identified based on mass spectral data associated with the chromatogram peaks.
[0046] Figure 4 is an exemplary implementation of steps 201 , 202 and 203 of the exemplary embodiment of the method set out in Figure 3.
[0047] Figure 5 is an exemplary embodiment of the method of Figure 1 where the sample used to obtain the chromatogram comprises derivatives of a known precursor.
[0048] Figure 6(a) is an exemplary implementation of steps 303a and 303b of the exemplary embodiment of the method set out in Figure 5 where the known precursor is cimifugin.
[0049] Figure 7(a) includes mass spectra of the known precursor, cimifugin, and its derivatives.
[0050] Figures 7(b) and 7(c) are zoomed in version of parts of the mass spectra of Figure 7(a).
[0051] Detailed Description of the Invention
[0052] Figure 1 shows a schematic arrangement of an apparatus that may be used for carrying out methods in accordance with embodiments of the present invention. The apparatus comprises a chromatography instrument 5 and a mass spectrometer 10 suitable for carrying out methods in accordance with embodiments of the present invention. The schematic diagram of the mass spectrometer 10 in Figure 1 represents the configuration of the Q-Exactive® mass spectrometer from Thermo Fisher Scientific, Inc.
[0053] Using the apparatus of Figure 1 , a sample to be analysed is supplied (for example from an autosampler) to the chromatography instrument 5. The chromatography instrument 5 is not shown in any detail in the figure but, in general, such a chromatography instrument 5 has a column and a detector (not shown). The sample is carried by a mobile phase through the column. The column has a stationary phase that may contain a plurality of particles. Compounds of the sample elute at different rates from the column according to their degree of interaction with the stationary phase. The compounds of the sample eluting from the column are detected by a detector.
[0054] A gas chromatography instrument employs a gas chromatography column and, in use, a sample is vapourised and injected into the gas chromatography column together with a carrier gas as the mobile phase. The carrier gas is typically an inert gas.
[0055] A liquid chromatography instrument employs a liquid chromatography column and, in use, a sample is dissolved in a liquid mobile phase and injected into the liquid chromatography column together with the liquid mobile phase. The mobile phase is pressurised in high performance liquid chromatography.
[0056] Detectors used in chromatography instruments are well known. Examples of liquid chromatography detectors include absorbance detectors, fluorescence detectors, evaporative light scattering detectors, fluorescence / chemiluminescence detectors, electrochemical and refractive index detectors. Examples of gas chromatography detectors include flame ionization detectors, thermal conductivity detectors, barrier discharge ionization detectors, electron capture detectors, sulphur chemiluminescence detectors, nitrogen phosphorus detectors, thermal conductivity detectors and pulsed discharge detectors.
[0057] A chromatogram may be produced by measuring the quantity of sample molecules which elute from the column over time using the detector. Sample components (compounds within the sample) which elute from the column will be detected as a peak above a baseline measurement on the chromatogram. Where different compounds within a sample have different elution rates, a plurality of peaks on the chromatograph may be detected. Ideally, individual sample peaks are separated in time from other peaks in the chromatogram such that different compounds do not interfere with each other.
[0058] Typically a chromatogram peak has a Gaussian shaped profile, or can be assumed to have a Gaussian shaped profile. Chromatography instruments and the resulting chromatograms produced are well known in the art and so will not be described in significant detail.
[0059] The chromatography instrument is coupled to a mass spectrometer and configured to provide the separated compounds to the mass spectrometer. These compounds are introduced, typically by injection, into the mass spectrometer. The mass spectrometer may be arranged to ionise (and optionally fragment) the injected compounds.
[0060] The mass spectrometer 130 is arranged to generate a mass spectrum of intensity (or abundance) against the mass-to-charge ratio (i.e. m / z value) of the ionised compounds (or fragments of those compounds). It is known that the generation of a mass spectrum may involve separation or selection of the ionised components according to their m / z value, followed by the measuring of a signal or signals caused by these separated groups of ions and / or ionised fragments. The separation or selection of the ionised components can happen in that ions of a specific m / z value have specific trajectories, on which these ions oscillate. Due to this oscillation, characteristic signals of the ions can be detected which have a frequency co and from which a specific m / z value can be assigned. The operation of mass spectrometers is well known in the art. A specific example of a mass spectrometer is an Orbitrap™ mass spectrometer, which is schematically shown in Figure 1 and described below. However, the skilled person would appreciate that the mass spectrometer may be of any type. For example the mass spectrometer may be any one of: a mass spectrometer comprising an ion trap (such as a linear ion trap spectrometer), a time of flight (TOF) mass spectrometer, a Fourier transform ion cyclotron resonance mass spectrometer (FT-ICRMS) or another type of electrostatic ion trap mass spectrometer.
[0061] The compounds separated by the chromatography instrument are typically received by the mass spectrometer as a function of the retention time. In this way it will be appreciated that the mass spectrometer receives compounds emitted at the same value (or within the same range of values) of the retention time simultaneously (or substantially simultaneously). Consequently, the mass spectrometer is arranged to generate mass spectra as a function of the retention time. Each mass spectrum may be thought of as (or representing) a mass spectrum of the compounds emitted at a respective value of the retention time. Each mass spectrum need not be a full mass spectrum in the sense of a complete intensity vs m / z plot across the entire m / z range. For example, a mass spectrum, as referred to herein, may comprise one or more m / z vs intensity data points. The mass spectrum may be limited to a particular m / z range of interest. The mass spectrum may comprise only the centroids within a particular m / z range of interest. The mass spectrometer, therefore, is arranged to produce mass spectral data for each respective value of the retention time. The mass spectral data for each respective value of the retention time may comprise one or more mass spectra.
[0062] Some devices, such as some time-of-flight mass spectrometers may not provide mass spectrometry data as an ordered set of plotted mass spectra with respect to the retention time. Instead the mass spectrometry data may be a stream of m / z-intensity value pairs, with associated values of the retention time.
[0063] A description of obtaining mass spectra in accordance with the exemplary mass spectrometer shown in Figure, which is the Q-Exactive® mass spectrometer from Thermo Fisher Scientific, Inc, is now provided.
[0064] The sample components separated via chromatography are ionized using an electrospray ionization source (ESI source) 20 which may be at atmospheric pressure. It will be appreciated by those skilled in the art that other suitable types of ionization source may be used, such as atmospheric pressure chemical ionization (APCI), thermospray ionization etc. Sample ions then enter a vacuum chamber of the mass spectrometer 10 and are directed by a capillary 25 into an RF-only S lens 30. The ions are focused by the S lens 30 into an injection flatpole 40 which injects the ions into a bent flatpole 50 with an axial field. The bent flatpole 50 guides (charged) ions along a curved path through it whilst unwanted neutral molecules such as entrained solvent molecules are not guided along the curved path and are lost.
[0065] An ion gate (TK lens) 60 is located at the distal end of the bent flatpole 50 and controls the passage of the ions from the bent flatpole 50 into a downstream quadrupole mass filter 70. The quadrupole mass filter 70 is typically but not necessarily segmented and, when operated in a selective mode, serves as a band pass filter, allowing passage of a selected mass to charge ratio or limited mass to charge ratio range whilst excluding ions of other mass to charge ratios (m / z).The quadrupole mass filter 70 can be operated to allow passage of ions of a relatively wide mass to charge ratio range (e.g. 400-1210 amu), in particular for MS1 scans, which is useful for acquiring a wide range mass spectrum.
[0066] Ions then pass through a quadrupole exit lens / split lens arrangement 80 and into a transfer multipole 90. The transfer multipole 90 guides the mass filtered ions from the quadrupole mass filter 70 into a curved trap (C-trap) 100. The C-trap 100 has longitudinally extending, curved electrodes which are supplied with RF voltages and end cap electrodes to which DC voltages are supplied to provide potential barriers at the ends of the C-trap 100. The result is a potential well that extends along the curved longitudinal axis of the C-trap 100. In a first mode of operation, the DC end cap voltages are set on the C-trap so that ions arriving from the transfer multipole 90 are captured in the potential well of the C-trap 100, where they are cooled. The injection time (IT) of the ions into the C-trap determines the number of ions (ion population) that is subsequently ejected from the C-trap into the mass analyser. Whilst a C-trap 100 is used in the mass spectrometer of Fig. 1 , in other embodiments, for example where a different type of mass analyser is used, different ion storage devices could be used instead, e.g. a linear trap with straight, not curved, electrodes.
[0067] Cooled ions reside in a cloud towards the bottom of the potential well and are then ejected orthogonally from the C-trap 100 towards an orbital trapping device 110 such as the Orbitrap® mass analyser sold by Thermo Fisher Scientific, Inc. The orbital trapping device 110 has an off centre injection aperture and the ions are injected into the orbital trapping device 110 as coherent packets, through the off centre injection aperture. Ions are then trapped within the orbital trapping device 110 by a hyperlogarithmic electric field, and undergo back and forth motion in a longitudinal direction whilst orbiting around the inner electrode.
[0068] The axial (z) component of the movement of the ion packets in the orbital trapping device 110 is (more or less) defined as simple harmonic motion, with the angular frequency in the z direction being related to the square root of the mass to charge ratio of a given ion species. Thus, over time, ions separate in accordance with their mass to charge ratio.
[0069] Ions in the orbital trapping device 110 are detected by use of an image detector (not shown in Figure 1) which produces a “transient” in the time domain containing information on all of the ion species as they pass the image detector. The transient is then subjected to a Fast Fourier Transform (FFT) resulting in a series of peaks in the frequency domain. From these peaks, a mass spectrum, representing abundance / ion intensity versus m / z, can be produced.
[0070] In the configuration described above, the sample ions (more specifically, a subset of the sample ions within a mass range of interest, selected by the quadrupole mass filter) are analysed by the orbital trapping device 110 without fragmentation. The resulting mass spectrum is denoted MS1 .
[0071] MS2 analysis (or, more generally, MSn) can also be carried out by the mass spectrometer 10 of Figure 1. To achieve this, precursor sample ions are generated and transported to the quadrupole mass filter 70 where a subsidiary mass range is selected. The ions that leave the quadrupole mass filter 70 are directed through the C-trap 100 to the fragmentation chamber 120. The fragmentation chamber 120 is, in the mass spectrometer 10 of Figure 1 , a higher energy collisional dissociation (HCD) device to which a collision gas is supplied. The potential applied to the fragmentation chamber 120 is such that precursor ions arriving into the fragmentation chamber 120 have sufficient energy that their collisions with collision gas molecules result in fragmentation of the precursor ions into fragment ions. The fragment ions are then ejected from the fragmentation chamber 120 back towards the C-trap 100, where they are once again trapped and cooled in the potential well. Finally, the fragment ions trapped in the C-trap are ejected orthogonally towards the orbital trapping device 110 for analysis and detection. The resulting mass spectrum of the fragment ions is denoted MS2.
[0072] Although an HCD fragmentation chamber 120 is shown in Figure 1 , other fragmentation devices may be employed instead, employing such methods as collision induced dissociation (CID), electron capture dissociation (ECD), electron transfer dissociation (ETD), photodissociation, and so forth.
[0073] The “dead end” configuration of the fragmentation chamber 120 in Figure 1 , wherein precursor ions are ejected axially from the C-trap 100 in a first direction towards the fragmentation chamber 120, and the resulting fragment ions are returned back to the C-trap 100 in the opposite direction, is described in further detail in WO-A-2006 / 103412.
[0074] The mass spectrometer 10 is under the control of a controller (not shown) which, for example, is configured to control the timing of ejection of the trapping components, to set the appropriate potentials on the electrodes of the quadrupole etc. so as to focus and filter the ions, to capture the mass spectral data from the orbital trapping device 110, control the sequence of MS1 and MS2 scans and so forth. It will be appreciated that the controller may comprises a computer that may be operated according to a computer program comprising instructions to cause the mass spectrometer to obtain mass spectra of the sample. Of course, in other embodiments, a mass spectrometer may comprise a fragmentation chamber arranged in a “fly through” configuration, where fragmented ions travel through the fragmentation chamber on to a further mass analyser, such as a linear ion trap, for MS2 analysis. For example, the Fusion Lumos Tribrid mass spectrometer from Thermo Fisher Scientific, Inc. comprises such a “fly through” fragmentation chamber.
[0075] It is to be understood that the specific arrangement of components shown in Figure 1 is not essential to the methods subsequently described. Figure 1 is only provided to give an example of an apparatus that could be used to obtain the chromatogram and associated mass spectral data analysed by the method of the claimed invention. Indeed, the method is not limited to actually performing chromatography and mass spectrometry to obtain the chromatogram and associated mass spectral data, since these could have been previously stored in a database from which they are obtained.
[0076] An embodiment of the method will now be described with reference to Figure 2. Steps 101 , 102
[0077] As set out in step 101 , first a chromatogram of the sample is obtained.
[0078] The sample comprises a mixture of a plurality of compounds, wherein at least two of the compounds are isomers. Isomers are compounds having the same molecular formula but a different structural formula. Optionally, the mixture of the plurality of compounds consists of (i.e. comprises only) isomers. The mixture of the plurality of compounds may comprise a first set of isomers having the same molecular formula but different structural formula and a second set of isomers having the same molecular formula but different structural formula where the molecular formula of the first set of isomers is different form the molecular formula of the second set of isomers.
[0079] The chromatogram may have been produced by gas chromatography or liquid chromatography of the sample. In one embodiment, the method may comprise the step of performing gas chromatography or liquid chromatography to obtain the chromatogram. Using the chromatography instrument 5 to obtain a chromatogram is set out above in the context of Figure 1.
[0080] Performing chromatography to obtain the chromatogram may comprise injecting a sample into the chromatography column of the chromatography instrument 5 where the sample is carried through the column by a mobile phase. The chromatography column may have a stationary phase that may contain a plurality of particles. Compounds of the sample elute at different rates from the column according to their degree of interaction with the stationary phase. The compounds of the sample eluting from the column are detected by a detector.
[0081] The chromatography instrument 5 may be a gas chromatography instrument or a liquid chromatography instrument.
[0082] A gas chromatography instrument employs a gas chromatography column and, in use, a sample is vapourised and injected into the gas chromatography column together with a carrier gas as the mobile phase. The carrier gas is typically an inert gas.
[0083] A liquid chromatography instrument employs a liquid chromatography column and, in use, a sample is dissolved in a liquid mobile phase and injected into the liquid chromatography column together with the liquid mobile phase. The mobile phase is pressurised in high performance liquid chromatography. The chromatogram is produced by measuring the quantity of sample molecules which elute from the column over time using the detector of the chromatography instrument. Sample components (compounds within the sample) which elute from the column will be detected as a peak above a baseline measurement on the chromatogram. Where different compounds within a sample have different elution rates, a plurality of peaks on the chromatograph may be detected. Ideally, individual sample peaks are separated in time from other peaks in the chromatogram such that different compounds do not interfere with each other.
[0084] A chromatogram is a plot of detector response (detector signal intensity) against retention time. The retention time of a compound is generally measured as the period of time between injection of sample into the column and the relative intensity peak maximum after chromatographic separation by the column. The retention time depends on the compound and also on the measurement conditions, such as the composition of the stationary phase, the pressure of the column, the temperature of the column and the dimensions of the column, as is known in the art.
[0085] Typically a chromatogram peak has a Gaussian shaped profile, or can be assumed to have a Gaussian shaped profile.
[0086] The method may alternatively not comprise the step of performing gas chromatography or liquid chromatography to obtain the chromatogram. Instead, the chromatogram of the sample may be pre-stored in and obtained from a memory or database.
[0087] Step 102 comprises obtaining mass spectral data associated with chromatogram peaks.
[0088] Steps 101 and 102 may be performed substantially simultaneously.
[0089] The associated mass spectral data for each peak may comprise one or more MSnmass spectra where n > 1 . The associated mass spectral data may be one of a plurality of mass spectra associated with the chromatogram peak. For example, each peak may have an associated MS1spectrum and MS2spectrum. Each peak may have a plurality of multi-stage MSnassociated mass spectra that may form a mass spectral tree. The mass spectral data may be a list of m / z peaks in the one or more MSnmass spectra where n > 1 . An MS1spectra of an ionised compound could contain some m / z peaks corresponding to fragments. For example, a MS1spectra obtained by GC / MS with electron impact analyzer typically would contain some m / z peaks corresponding to fragments.
[0090] The mass spectral data may be obtained using the mass spectrometer 10 coupled to the chromatography instrument 5 as discussed in the exemplary context of Figure 1 . The compounds separated by the chromatography instrument 5 may be introduced into the mass spectrometer 10 according to their retention time. Once introduced into the mass spectrometer 10, the compounds will be ionised, optionally fragmented and separated according to their m / z value. The detector (not shown) of the mass spectrometer 10 may measure signals generated by the ionised sample compounds / fragments to generate mass spectra as a function of retention time. Each mass spectrum generated is associated with a chromatogram peak corresponding to a certain retention time.
[0091] The associated mass spectral data for each peak may comprise one or more MSnmass spectra where n > 1 . The associated mass spectral data may be one of a plurality of mass spectra associated with the chromatogram peak. For example, each peak may have an associated MS1spectrum and MS2spectrum. Each peak may have a plurality of multi-stage MSnassociated mass spectra that may form a mass spectral tree. The mass spectral data may be a list of m / z peaks in the one or more MSnmass spectra where n > 1 .
[0092] The method may alternatively not comprise the step of performing mass spectrometry to obtain the mass spectral data. Instead, the chromatogram of the sample and associated mass spectral data may be pre-stored in a memory or database.
[0093] Step 103
[0094] Step 103 comprises identifying fragments of each candidate chemical structure of a first plurality of candidate chemical structures and identifying a m / z ratio of at least one unique fragment for each candidate chemical structure.
[0095] Candidate chemical structures are compounds that are predicted to be compounds present in the sample and so correspond to peaks of the chromatogram. The candidate chemical structures are defined according to their structural formula. The candidate chemical structures may also be referred to as candidate compounds or candidate structural formulas. The candidate chemical structures may be defined as a 2D representation of the molecular structure, as shown in Figures 4 and 6. Candidate isomers are candidate chemical structures that are also isomers i.e. having the same molecular formula but a different structural formula. The candidate isomers may be defined according to their structural formulas.
[0096] The first plurality of candidate chemical structures comprises candidate isomers. Optionally, the first plurality of candidate chemical structures consist of (comprise only) candidate isomers. A plurality of possible fragments of the candidate chemical structures are predicted by performing fragmentation of the candidate chemical structures in-silico (i.e. by computer modelling or simulation). Optionally, all possible fragments of the candidate chemical structures of the first plurality of candidate chemical structures are predicted by performing fragmentation of the candidate chemical structures in-silico. Performing such fragmentation in-silico based on the structural formula of the compound is well known in the art and may be performed using, for example, Mass Frontier. The fragments generated in-silico for one candidate chemical structure that are not generated in-silico for the other candidate chemical structure(s) within the first plurality of candidate chemical structures are the unique fragments. The m / z ratios of the unique fragment(s) for each candidate chemical structure are then obtained, for example, using Mass Frontier. The unique fragment of a candidate chemical structure has a m / z ratio that is different from the m / z ratios of the possible fragments predicted for the other candidate chemical structure within the first plurality of candidate chemical structures. A plurality of unique fragments may be predicted for a candidate chemical structure. The fragments predicted in-silico may not all be produced during actual (i.e. experimental rather than simulated) fragmentation of the candidate isomer. Consequently, some of the unique fragments predicted for a candidate chemical structure may not be produced during actual (i.e. experimental rather than simulated) fragmentation and so peaks at m / z ratios corresponding to the m / z ratio of those unique fragments would not appear in the mass spectrum of the candidate chemical structure, as explained in Examples 1 ,2, 3 and 4. However, the simulated fragmentation techniques do predict more correct fragments than incorrect fragments for a compound and so the method of the claimed invention still enables assignment of candidate chemical structures, particular candidate chemical structures that are isomers, to the correct chromatogram peaks according to the method described herein, as exemplified by Examples 1 , 2, 3 and 4.
[0097] The first plurality of candidate chemical structures may form part of a list of candidate chemical structures and the method may comprise the step of obtaining a list of candidate chemical structures. The list of candidate chemical structures may comprise additional candidate chemical structures to the first plurality of candidate chemical structures. For example, the list of candidate chemical structures may comprise a second plurality of candidate chemical structures in addition to the first plurality of candidate chemical structures. The list of candidate chemical structures may consist of the first plurality of candidate chemical structures.
[0098] The list of candidate chemical structures may be obtained based on (i) the associated mass spectral data of each chromatogram peak and / or based on (ii) knowledge of a precursor to one or more compound(s) of the plurality of compounds or knowledge of a molecular formula of one or more compound(s) of the plurality of compounds. An exemplary embodiment of the method according to option (i) is set out in Figures 3 and 4 an exemplary embodiment of the method according to option (ii) knowledge of a precursor to one or more compound(s) is set out in Figures 5 and 6. These embodiments are discussed in the context of those figures below. An exemplary embodiment of the method may comprise obtaining some candidate compounds of the list of the candidate chemical structures and assigning one or more of those candidate chemical structures to their most probable chromatogram peaks according to option (i) as discussed in Figure 3 and 4 and obtaining some candidate compounds of the list of the candidate chemical structures and assigning one or more of those candidate chemical structures to their most probable chromatogram peaks according to option (ii) as discussed in Figures 5 and 6.
[0099] In the embodiment according to option (ii), step 103 may be performed before or after or simultaneously with steps 101 and 102.
[0100] In the exemplary embodiment where there is knowledge of the molecular formula of the sample containing a mixture of isomers but no knowledge of the structural formulas of the isomers, the candidate isomers could be determined based on knowledge of their molecular formula. For example, a database or look-up table may be used, such as ChemSpider, PubChem, mzCloud, where based on the molecular formula, a list of all possible structural formulas and so all possible isomers is obtained.
[0101] Step 104
[0102] Step 104 comprises, for one or more of the candidate chemical structures of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak based on identifying one or more m / z peak(s) in the associated mass spectral data at m / z ratios corresponding to the m / z ratio(s) of the at least one unique fragment(s) of the candidate chemical structure. In this step, the candidate chemical structure is assigned to the most probable chromatogram peak. At least one of the candidate chemical structure(s) assigned to a chromatogram peak is one of the candidate isomers.
[0103] This step of assigning the candidate chemical structure of the first plurality of candidate chemical structures to a chromatogram peak of the plurality of chromatogram peaks may be performed for each of the candidate isomers so that each candidate isomer is assigned to a chromatogram peak (the most probable chromatogram peak for that candidate isomer). The most probable chromatogram peak is the chromatogram peak within the plurality of chromatogram peaks that is most likely to correspond to the candidate chemical structure. Optionally, the mixture of the plurality of compounds consists of isomers (i.e. contains only isomers). Further optionally, step 104 of the method may comprise assigning multiple candidate chemical structures that are candidate isomers to a chromatogram peak such that each chromatogram peak has a single candidate isomer assigned thereto.
[0104] The step of assigning the candidate chemical structure to a chromatogram peak may comprise determining a score for each chromatogram peak of the plurality of chromatogram peaks indicating the extent to which the mass spectral data associated with the chromatogram peak correlates / corresponds to the candidate chemical structure and assigning the candidate chemical structure to the chromatogram peak having the score indicating the greatest correlation. In other words, the score indicates the likelihood of the chromatogram peak corresponding to the candidate chemical structure and the candidate chemical structure is assigned to the chromatogram peak having the score indicative of the highest likelihood. The score indicative of the highest likelihood may be the highest score.
[0105] The score may be calculated based on identifying one or more m / z peak(s) in the associated mass spectral data that correspond to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure.
[0106] In other words, the step of assigning the candidate chemical structure to a chromatogram peak may comprise determining a score indicating the number and / or intensity of one or more m / z peak(s) at m / z ratios corresponding to the m / z ratio(s) of the least one unique fragment(s) of the candidate chemical structure in the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak having the score indicating the greatest number and / or intensity of one or more m / z peak(s) at m / z ratios corresponding to the m / z ratio(s) of the at least one unique fragment(s) of the candidate chemical structure in the associated mass spectral data. By “corresponding” in this context, this refers to the m / z peaks of the associated mass spectral data at m / z ratios that are the same as or within a tolerance range of the m / z ratio(s) of the at least one unique fragment(s) of the candidate chemical structure.
[0107] As discussed above, the associated mass spectral data for each peak may comprise one or more MSnmass spectra where n > 1 . The associated mass spectral data may be one of a plurality of mass spectra associated with the chromatogram peak. For example, each peak may have an associated MS1spectrum and MS2spectrum. Each peak may have a plurality of multistage MSnassociated mass spectra that may form a mass spectral tree. The mass spectral data may be a list of m / z peaks identified in the one or more MSnmass spectra where n > 1 . Each score may be calculated based on identifying a number and / or intensity of one or more m / z peak(s) corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure in the associated mass spectral data. The score value indicates the extent to which the candidate chemical structure corresponds to the mass spectral data associated with the chromatogram peak. The greater the number and / or intensity of the one or more m / z peak(s) corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure in the associated mass spectral data, the greater the correlation. By “corresponding” in this context, this refers to the m / z peaks of the associated mass spectral data at m / z ratios that are the same as or within a tolerance range of the m / z ratio(s) of the at least one unique fragment(s) of the candidate chemical structure.
[0108] The score may be calculated based on, or may be, a ratio of intensities of the one or more m / z peak(s) corresponding to the m / z ratio(s) of unique fragment(s) of the candidate chemical structure to intensities of all m / z peaks in the associated mass spectral data. This may be referred to as an intensity ratio. The score may be calculated based on, or may be, the absolute ratio of intensities or the relative ratio of intensities. The relative ratio of intensities may be referred to as the relative intensity ratio. The absolute ratio of intensities may be referred to as the absolute intensity ratio. The relative intensity ratios may be normalised values of the absolute intensity ratios. In this example, the candidate chemical structure is assigned to the chromatogram peak with the associated mass spectral data for which the ratio is greatest.
[0109] The score may be calculated based on, or may be, a ratio of a total number of the m / z peak(s) that are identified in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio(s) of unique fragment(s) of the candidate chemical structure to a number of all m / z peaks in the associated mass spectral data of the chromatogram peak. This ratio may be referred to as an absolute peak count ratio. The absolute peak count ratios may be normalised and such normalised ratios may be referred to as relative peak count ratios.
[0110] In embodiments where the score is based on the absolute peak count ratio, the candidate chemical structure may be assigned to the chromatogram peak with associated mass spectral data for which the absolute peak count ratio is greatest.
[0111] In embodiments where the score is based on the relative (normalised) peak count ratio, the candidate chemical structure may be assigned to the chromatogram peak with associated mass spectral data for which the relative peak count ratio is greatest. The score may be a value calculated based on a combination of the ratios above. For example, the score may be a value calculated based on both the intensity ratio (the ratio of intensities of the one or more m / z peak(s) that are identified in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio(s) of unique fragment(s) of the candidate chemical structure to intensities of all m / z peaks in the associated mass spectral data of the chromatogram peak) and the peak count ratio (the ratio of a number of the one or more m / z peak(s) that are identified in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio(s) of unique fragment(s) of the candidate chemical structure to a number of all m / z peaks in the associated mass spectral data of the chromatogram peak).
[0112] The score may be a sum of the intensity ratio and the peak count ratio. The candidate chemical structure may be assigned to the chromatogram peak with associated mass spectral data having the highest score, which is the sum of both the intensity ratio and the peak count ratio. The score may be a weighted sum or average of both the intensity ratio and the peak count ratio.
[0113] For the candidate chemical structure, both a score calculated based on the intensity ratio and a score calculated based on the peak count ratio may be calculated for the candidate chemical structure in respect of the mass spectral data of each chromatogram peak. Where a candidate chemical structure has different scores of the same type (e.g. intensity or peak count based) in respect of the associated mass spectral data of different chromatogram peaks, those scores may be used to assign the candidate chemical structure to the corresponding chromatogram peak. The difference between the scores may be above a threshold for those scores to be used to assign the candidate chemical structure to the most likely chromatogram peak. For example, for a candidate chemical structure, if the score based on the intensity ratio for the mass spectral data of one chromatogram peak is similar to the score based on the intensity ratio for the mass spectral data of the other chromatogram peak but the score based on the peak count ratio for the mass spectral data of one chromatogram peak is different from the score based on the peak count ratio for the mass spectral data of the other chromatogram peak, then the scores based on the peak count ratio may be used to assign the candidate chemical structure to the most probable chromatogram peak. The intensity ratio may be the absolute intensity ratio or the relative intensity ratio. The peak count ratio may be the absolute peak count ratio or the relative peak count ratio. The step of assigning the candidate chemical structure to a chromatogram peak may be performed for each of the candidate chemical structures such that the method comprises determining scores for each of the candidate chemical structures, each score indicating the extent to which the mass spectral data associated with the chromatogram peak correlates / corresponds to that candidate chemical structure and assigning each candidate chemical structure to the chromatogram peak having the score indicating the greatest correlation.
[0114] The scores may be represented in a matrix. Each cell of the table may correspond to the score value indicating the extent to which the candidate chemical structure corresponds to the mass spectral data associated with the chromatogram peak. The candidate chemical structures may be set out in the rows and the associated mass spectral data for each chromatogram peak set out in the columns or vice versa.
[0115] Worked Examples for Calculating Scores
[0116] Examples of calculating scores for candidate chemical structures that are candidate isomers to one of two chromatogram peaks are set out below. These examples demonstrate how the scores may be calculated based on identifying one or more m / z peak(s) in a mass spectrum at m / z ratios corresponding to the unique fragment(s) of the candidate isomer. In these examples, the mass spectrum is a MS1mass spectrum. To verify that the scoring system works to correlate candidate isomers to the correct mass spectrum (and so the correct chromatogram peak associated with that mass spectrum), these examples use a mass spectrum of each of the compound (specifically each candidate isomer) obtained experimentally. For these examples, the mass spectrum of each compound was obtained using a Thermo Fisher GC / MS Orbitrap instrument employing an electron impact analyser. These scores are based on identifying one or more m / z peak(s) in the mass spectrum of a chromatogram peak at m / z ratio(s) that correspond to the m / z of the unique fragments of the candidate isomer. The skilled person would understand how this scoring system could be applied to assign the candidate isomer to the corresponding mass spectrum associated with a chromatogram peak where it is not already known which mass spectrum corresponds to which candidate isomer.
[0117] Example 1
[0118] Example 1 sets out exemplary ways of calculating scores for assigning candidate isomers to the correct mass spectrum. In this example, the candidate isomers are o-anisidine and m-anisidine. A mass spectrum for o-anisidine and a mass spectrum for m-anisidine are experimentally obtained using a Thermo Fisher GC / MS Orbitrap instrument. The candidate isomers are assigned to the correct experimental mass spectrum based on generating a score indicating the extent to which the candidate isomer correlates to the experimental mass spectrum. The score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratio(s) that correspond to the m / z of the unique fragment(s) of the candidate isomer.
[0119] Table 1 : Candidate Isomers Structural Formula Data
[0120] The predicted fragments can be determined by performing fragmentation of the compound in- silico (i.e. by computer modelling or simulation). Performing such fragmentation in-silico based on the structural formula of the compound is well known in the art and may be performed using, for example, Mass Frontier. The m / z values of the fragments generated in-silico for that isomer that have are different compared to the m / z values of fragments generated for the other isomer are the unique fragments (the fragments are non-isobaric). In this case, o-anisidine has 7 unique fragments (fragments generated for o-anisidine but not generated for m-anisidine) and m- anisidine has 4 unique fragments (fragments generated for m-anisidine but not generated for o- anisidine).
[0121] Table 2:
[0122] Total number of m / z peaks in the experimental mass spectra corresponding to the m / z ratios of unique fragments of the candidate isomers The experimental mass spectra are analysed to determine whether there are any m / z peaks at m / z corresponding to the unique fragments of each isomer. In this case, the o-anisidine experimental mass spectrum has two m / z peaks at m / z ratios corresponding to the unique fragments of o-anisidine. The o-anisidine experimental mass spectrum also has two m / z peaks at m / z ratios corresponding to the unique fragments of m-anisidine. The m-anisidine experimental mass spectrum has one m / z peak at a m / z ratio corresponding to a unique fragment of o-anisidine. The m-anisidine experimental mass spectrum has two m / z peaks at m / z ratios corresponding to the unique fragments of m-anisidine.
[0123] Ideally, there would be a non-zero value for number of m / z peaks corresponding to unique fragments of the first isomer in the corresponding experimental mass spectra and a zero value for number of m / z peaks corresponding to unique fragments of the first isomer in the noncorresponding mass spectrum. This is because each mass spectrum should only correspond to one of the isomers and so m / z peaks corresponding to unique fragments of that isomer should only be present in the mass spectrum of that isomer.
[0124] In this case, both of the experimental mass spectra include m / z peaks at m / z corresponding to unique fragments for both isomers. This means that the in-silico fragmentation predicted a fragment for only one of the isomers and so identified this fragment as a unique fragment but such a fragment was generated for both isomers in practice (i.e. during experimental fragmentation). This is therefore a false negative prediction for that fragment. However, as evidenced by these examples, despite erroneous identification of a fragment as being unique for that isomer, the fragmentation model predicts correct fragments with a higher probability than incorrect ones, and the resulting score can consequently be used to correctly assign isomers to the corresponding experimental mass spectra and so the corresponding chromatogram peaks.
[0125] Despite erroneously identifying a fragment as a unique fragment when it is generated during fragmentation of both isomers, it will be appreciated from these examples that the in-silico fragmentation predicts correct fragments with a higher probability than incorrect ones, and the resulting score can consequently be used to correctly assign isomers to the corresponding mass spectra and so corresponding chromatogram peaks.
[0126] The number of m / z peaks corresponding to the m / z of unique fragments of the isomers in the experimental mass spectra set out in Table 2 are used to generate the scores set out in Tables 3 to 5.
[0127] Table 3: Absolute Intensity Ratio Scores
[0128] The scores set out in Table 3 are the absolute ratio of intensities of m / z peak(s) at m / z ratios corresponding to the unique fragments of the candidate isomer to the sum of the absolute intensities of all m / z peaks in each experimental mass spectrum.
[0129] These scores have a low value because the number and intensity of the m / z peaks at m / z ratios corresponding to unique fragments is low compared to the intensities of all peaks. The scores can be compared. The scores based on identifying m / z peaks at m / z ratios corresponding to unique fragments of m-anisidine in the o-anisidine experimental mass spectrum and in the m- anisidine experimental mass spectrum are the same and so cannot be used to determine which chromatogram peak corresponds to m-anisidine. However, the scores based on identifying m / z peaks at m / z ratios corresponding to unique fragments of o-anisidine are different for each experimental mass spectrum. The o-anisidine experimental mass spectrum has a higher intensity ratio for candidate isomer o-anisidine (the ratio of intensities of m / z peak(s) at m / z ratios corresponding to the unique fragments of the o-anisidine to intensities of all m / z peaks in the o- anisidine experimental mass spectrum) compared to the m-anisidine experimental mass spectrum (the ratio of intensities of m / z peak(s) at m / z ratios corresponding to the unique fragments of o-anisidine to intensities of all m / z peaks in the m-anisidine experimental mass spectrum). Therefore, o-anisidine can be correctly assigned to the o-anisidine experimental mass spectrum. By the process of elimination, m-anisidine can be correctly assigned to the m-anisidine experimental mass spectrum.
[0130] Table 4: Relative Intensity Ratio Scores
[0131] The scores set out in Table 4 are the relative (normalised) ratio of intensities of m / z peak(s) at m / z ratios corresponding to the unique fragments of the candidate isomer to intensities of all m / z peaks in each experimental mass spectrum. The relative intensity ratio is calculated by subtracting the lowest absolute intensity ratio from the absolute intensity ratio and dividing by the highest absolute intensity ratio. By way of example, the o-anisidine experimental mass spectrum has a relative intensity ratio for o-anisidine calculated by: (0.000323207-0.00020972) I 0.035898978 .
[0132] As discussed above in respect of Table 3, similarly, the scores in Table 4 based on identifying m / z peaks at m / z ratios corresponding to unique fragments of o-anisidine are different for the o- anisidine experimental mass spectrum compared to the m-anisidine experimental mass spectrum. Therefore, o-anisidine can be correctly assigned to the o-anisidine experimental mass spectrum based on the higher score compared to the score for the m-anisidine experimental mass spectrum. In other words, o-anisidine can be correctly assigned to the corresponding experimental mass spectrum based on the highest scoring experimental mass spectrum (highest score in that row).
[0133] Table 5: Absolute Peak Count Ratio Scores
[0134] The scores set out in Table 5 are the absolute ratio of the number of m / z peak(s) corresponding to m / z ratios of the unique fragments of the candidate isomer to the number of all m / z peaks in each experimental mass spectrum. This ratio is referred to herein as the absolute peak count ratio.
[0135] The scores based on identifying m / z peaks at m / z ratios corresponding to unique fragments of m- anisidine in each experimental mass spectrum are the same and so cannot be used to determine which experimental mass spectrum corresponds to m-anisidine. However, the scores based on identifying m / z peaks at m / z ratios correspond to unique fragments of o-anisidine are different for each experimental mass spectrum. The o-anisidine experimental mass spectrum has a higher peak count ratio for o-anisidine (the ratio of number of m / z peak(s) at m / z ratios corresponding to unique fragments of o-anisidine to number of all m / z peaks in the o-anisidine experimental mass spectrum) compared to the peak count ratio of the m-anisidine experimental mass spectrum for o-anisidine (the ratio of number of m / z peak(s) at m / z ratios corresponding to unique fragments of o-anisidine to the number of all m / z peaks in the m-anisidine experimental mass spectrum). Therefore, o-anisidine can be correctly assigned to the o-anisidine experimental mass spectrum. By the process of elimination, m-anisidine can be correctly assigned to the m-anisidine experimental mass spectrum. Example 2
[0136] Example 2 sets out exemplary ways of calculating scores for assigning candidate isomers to the correct mass spectrum similarly to Example 1. In this example, the candidate isomers are 3- dimethylamino benzoic acid and 4-dimethylamino benzoic acid. A mass spectrum for 3- dimethylamino benzoic acid and a mass spectrum for 4-dimethylamino benzoic acid are experimentally obtained using a Thermo Fisher GC / MS Orbitrap instrument. The candidate isomers are assigned to the correct experimental mass spectrum based on generating a score indicating the extent to which the candidate isomer correlates to the experimental mass spectrum. The score is calculated based on identifying m / z peaks in the associated mass spectral data that correspond to the m / z of the unique fragment(s) of the candidate isomer.
[0137] Table 1 : Candidate Isomers Structural Formula Data
[0138] The total number of fragments can be predicted by performing fragmentation of the compound in- silico (i.e. by computer modelling or simulation). Performing such fragmentation in-silico based on the structural formula of the compound is well known in the art and may be performed using, for example, Mass Frontier. The fragments generated in-silico for that isomer that are not generated for the other isomer are the unique fragments. In this case, 3-dimethylamino benzoic acid has 23 unique fragments (fragments generated for 3-dimethylamino benzoic acid but not generated for 4-dimethylamino benzoic acid) and 4-dimethylamino benzoic acid has 33 unique fragments (fragments generated for 4-dimethylamino benzoic acid but not generated for 3- dimethylamino benzoic acid).
[0139] Table 2: Total number of m / z peaks in the experimental mass spectra corresponding to the m / z ratios of unique fragments of the candidate isomers
[0140] The 3-dimethylamino benzoic acid experimental mass spectrum is analysed to determine whether there are any m / z peaks at m / z ratios corresponding to the unique fragments of each isomer. In this case, the experimental mass spectrum for 3-dimethylamino benzoic acid has eleven m / z peaks at m / z ratios corresponding to the unique fragments of 3-dimethylamino benzoic acid. The 3-dimethylamino benzoic acid experimental mass spectrum also has thirteen m / z peaks at m / z ratios corresponding to the unique fragments of 4-dimethylamino benzoic acid. The 4-dimethylamino benzoic acid experimental mass spectrum has eight m / z peak at m / z ratios corresponding to unique fragments of 4-dimethylamino benzoic acid. The 4-dimethylamino benzoic acid experimental mass spectrum has twelve m / z peaks at m / z ratios corresponding to the unique fragments of 4-dimethylamino benzoic acid.
[0141] Ideally, there would be a non-zero value for number of m / z peaks corresponding to unique fragments of the first isomer in the corresponding experimental mass spectra and a zero value for number of m / z peaks corresponding to unique fragments of the first isomer in the noncorresponding mass spectrum. This is because each mass spectrum should only correspond to one of the isomers and so m / z peaks corresponding to unique fragments of that isomer should only be present in the mass spectrum of that isomer.
[0142] In this case, both of the experimental mass spectra include m / z peaks at m / z ratios corresponding to unique fragments for both isomers. This means that the in-silico fragmentation predicted a fragment for only one of the isomers and so identified this fragment as a unique fragment but such a fragment was generated for both isomers in practice (i.e. during experimental fragmentation). This is therefore a false negative prediction for that fragment. However, as discussed above, despite erroneous identification of a fragment as being unique for that isomer, the fragmentation model predicts correct fragments with a higher probability than incorrect ones, and the resulting score can consequently be used to correctly assign isomers to the corresponding mass spectrum, and so corresponding chromatogram peak. The number of m / z peaks at m / z ratios corresponding to the m / z ratios of unique fragments of the isomers in the experimental mass spectra as set out in Table 2 are used to generate the scores set out in Tables 3 to 5 of Example 2.
[0143] Table 3: Absolute Intensity Ratio Scores
[0144] The scores set out in Table 3 are the absolute intensity ratios of intensities of m / z peak(s) at m / z corresponding to the unique fragments of the candidate isomer to the sum of the absolute intensities of all m / z peaks in each experimental mass spectrum.
[0145] The 4-dimethylamino benzoic acid experimental mass spectrum has a higher absolute intensity ratio for 4-dimenthylamino benzoic acid than the 3-dimethylamino benzoic acid experimental mass spectrum. The 3-dimethylamino benzoic acid experimental mass spectrum has a higher absolute intensity ratio for 3-dimenthylamino benzoic acid than the 4-dimethylamino benzoic acid experimental mass spectrum.
[0146] Therefore, these absolute intensity ratio scores correctly identify 4-dimethylamino benzoic acid as corresponding to the experimental mass spectrum for 4-dimethylamino benzoic acid and 3- dimethylamino benzoic acid as corresponding to the experimental mass spectrum for 3- dimethylamino benzoic acid.
[0147] Table 4: Relative Intensity Ratio Scores
[0148] The scores set out in Table 4 are the relative ratio of intensities of m / z peak(s) at m / z ratios corresponding to the unique fragments of the candidate isomer to intensities of all m / z peaks in each experimental mass spectrum. These ratios are the normalised ratio of intensities of m / z peak(s) at m / z ratios corresponding to the unique fragments of the candidate isomer to intensities of all m / z peaks in each experimental mass spectrum.
[0149] The 4-dimethylamino benzoic acid experimental mass spectrum has a higher relative intensity ratio for 4-dimenthylamino benzoic acid than the 3-dimethylamino benzoic acid experimental mass spectrum. The 3-dimethylamino benzoic acid experimental mass spectrum has a higher relative intensity ratio for 3-dimenthylamino benzoic acid than the 4-dimethylamino benzoic acid experimental mass spectrum.
[0150] Therefore, these scores correctly identify 4-dimethylamino benzoic acid as corresponding to the experimental mass spectrum for 4-dimethylamino benzoic acid and 3-dimethylamino benzoic acid as corresponding to the experimental mass spectrum for 3-dimethylamino benzoic acid.
[0151] Table 5: Absolute Peak Count Ratio Score
[0152] The scores set out in Table 5 are the absolute peak count ratios. These absolute peak count ratios are the ratios of the number of m / z peak(s) corresponding to m / z ratios of the unique fragments of the candidate isomer to the number of all m / z peaks in each experimental mass spectrum.
[0153] The 4-dimethylamino benzoic acid experimental mass spectrum has a higher absolute peak count ratio for the 4-dimethylamino benzoic acid than the 3-dimethylamino benzoic acid experimental mass spectrum. The 3-dimethylamino benzoic acid experimental mass spectrum has a higher absolute peak count ratio for 3-dimethylamino benzoic acid than the 4- dimethylamino benzoic acid experimental mass spectrum.
[0154] Therefore, these absolute peak count ratio scores correctly identify 4-dimethylamino benzoic acid as correctly corresponding to the 4-dimethylamino benzoic acid experimental mass spectrum and 3-dimethylamino benzoic acid as correctly corresponding to the 3-dimethylamino benzoic acid experimental mass spectrum. As discussed above, for these scores, which are the relative ratio of a number of m / z peak(s) corresponding to the m / z of unique fragments of the candidate isomer to a number of all m / z peaks in the experimental mass spectrum, the combination of the candidate isomer and the corresponding experimental mass spectrum should have a value of one. The combination of the candidate isomer and the non-corresponding experimental mass spectrum should have a value lower than one but a value equal to one indicates the impossibility of a decision.
[0155] In this case, the score being less than one for the 4-dimethylamino benzoic acid experimental mass spectrum and candidate isomer 3-dimethylamino benzoic acid indicates that candidate isomer 3-dimethylamino benzoic acid does not correspond with the 4-dimethylamino benzoic acid experimental mass spectrum and so, by the process of elimination, corresponds with the 3- dimethylamino benzoic acid experimental mass spectrum. By the process of elimination, it can be correctly determined that the 4-dimethylamino benzoic acid candidate isomer corresponds with the 4-dimethylamino benzoic acid experimental mass spectrum.
[0156] Example 3
[0157] Example 3 sets out exemplary ways of calculating scores for assigning candidate isomers to the correct mass spectrum similarly to Examples 1 and 2. In this example, the candidate isomers are 3-methylamino benzoic acid and 4-methylamino benzoic acid. A mass spectrum for 3- methylamino benzoic acid and a mass spectrum for 4-methylamino benzoic acid are experimentally obtained using a Thermo Fisher GC / MS Orbitrap instrument. The candidate isomers are assigned to the correct experimental mass spectrum based on generating a score indicating the extent to which the candidate isomer correlates to the experimental mass spectrum. The score is calculated based on identifying m / z peaks in the associated mass spectral data that correspond to the m / z of the unique fragment(s) of the candidate isomer.
[0158] Table 1 Candidate Isomers Structural Formula Data
[0159] The total number of fragments can be determined by performing fragmentation of the compound in-silico (i.e. by computer modelling or simulation). Performing such fragmentation in-silico based on the structural formula of the compound is well known in the art and may be performed using, for example, Mass Frontier. The fragments generated in-silico for that isomer that are not generated fort the other isomer are the unique fragments. In this case, 3-methylamino benzoic acid has 21 unique fragments (fragments generated for 3-methylamino benzoic acid but not generated for 4-methylamino benzoic acid) and 4-methylamino benzoic acid has 27 unique fragments (fragments generated for 4-methylamino benzoic acid but not generated for 3- methylamino benzoic acid).
[0160] Table 2: Total number of all m / z peaks in the experimental mass spectra corresponding to the m / z ratios of unique fragments of the candidate isomers
[0161] The 3-methylamino benzoic acid experimental mass spectrum is analysed to determine whether there are any m / z peaks at m / z corresponding to the unique fragments of each isomer. In this case, the 3-methylamino benzoic acid experimental mass spectrum has ten m / z peaks at m / z ratios corresponding to the unique fragments of 3-methylamino benzoic acid. The 3- methylamino benzoic acid experimental mass spectrum also has nine m / z peaks at m / z ratios corresponding to the unique fragments of 4-methylamino benzoic acid. The 4-methylamino benzoic acid experimental mass spectrum has ten m / z peaks at m / z ratios corresponding to unique fragments of 3-methylamino benzoic acid. The 4-methylamino benzoic acid experimental mass spectrum has eight m / z peaks at m / z ratios corresponding to the unique fragments of 4- methylamino benzoic acid. Ideally, there would be a non-zero value for number of m / z peaks corresponding to unique fragments of the first isomer in the corresponding experimental mass spectra and a zero value for number of m / z peaks corresponding to unique fragments of the first isomer in the noncorresponding mass spectrum. This is because each mass spectrum should only correspond to one of the isomers and so m / z peaks corresponding to unique fragments of that isomer should only be present in the mass spectrum of that isomer.
[0162] In this case, both of the experimental mass spectra include m / z peaks at m / z ratios corresponding to unique fragments for both isomers. This means that the in-silico fragmentation predicted a fragment for only one of the isomers and so identified this fragment as a unique fragment but such a fragment was generated for both isomers in practice (i.e. during experimental fragmentation). This is therefore a false negative prediction for that fragment.
[0163] However, as discussed above, despite erroneous identification of a fragment as being unique for that isomer, the fragmentation model predicts correct fragments with a higher probability than incorrect ones, and the resulting score can consequently be used to correctly assign isomers to the corresponding mass spectrum, and so corresponding chromatogram peak.
[0164] The number of m / z peaks corresponding to the m / z of unique fragments of the isomers in the experimental mass spectra as set out in Table 2 are used to generate the scores set out in Tables 3 to 5.
[0165] Table 3: Absolute Intensity Ratio Score
[0166] The scores set out in Table 3 are the absolute intensity ratios of intensities of m / z peak(s) at m / z corresponding to the unique fragments of the candidate isomer to the sum of the absolute intensities of all m / z peaks in each experimental mass spectrum.
[0167] The 3-methylamino benzoic acid experimental mass spectrum has a higher absolute intensity ratio for 3-methylamine benzoic acid than the experimental mass spectrum for 4-methylamino benzoic acid. The 4-methylamino benzoic acid has a higher absolute intensity ratio for 4-methylamine benzoic acid than the experimental mass spectrum for 3-methylamino benzoic acid.
[0168] Therefore, these scores identify 3-methylamino benzoic acid as correctly corresponding to the experimental mass spectrum for 3-methylamine benzoic acid and 4-dimethylamino benzoic acid as correctly corresponding to the experimental mass spectrum for 4-methylamine benzoic acid
[0169] Table 4: Relative Intensity Ratio Score
[0170] The scores set out in Table 4 are the relative (normalised) ratio of intensities of m / z peak(s) at m / z ratios corresponding to the unique fragments of the candidate isomer to intensities of all m / z peaks in each experimental mass spectrum.
[0171] The 3-methylamino benzoic acid experimental mass spectrum has a higher relative intensity ratio for candidate isomer 3-methylamino benzoic acid than the 4-methylamino benzoic acid experimental mass spectrum.
[0172] The 4-methylamino benzoic acid experimental mass spectrum has a higher relative intensity ratio for candidate isomer 4-methylamino benzoic acid than the 3-methylamino benzoic acid experimental mass spectrum.
[0173] Therefore, these relative intensity ratios correctly identify 3-methylamino benzoic acid as corresponding to the 3-methylamino benzoic acid experimental mass spectrum and 4- methylamino benzoic acid as correctly corresponding to the 4-methylamino benzoic acid experimental mass spectrum.
[0174] Table 5: Absolute Peak Count Ratio Score
[0175] The scores set out in Table 5 are the absolute peak count ratios. The absolute peak count ratios are the ratios of the number of m / z peak(s) corresponding to m / z ratios of the unique fragments of the candidate isomer to the number of all m / z peaks in each experimental mass spectrum.
[0176] The 3-methylamino benzoic acid experimental mass spectrum has a higher absolute peak number ratio for candidate isomer 3-methylamino benzoic acid than the 4-methylamino benzoic acid experimental mass spectrum.
[0177] The 4-methylamino benzoic acid experimental mass spectrum has a higher absolute peak count ratio for candidate isomer 4-methylamino benzoic acid than the 3-methylamino benzoic acid experimental mass spectrum.
[0178] Therefore, these absolute peak count ratios correctly identify 3-methylamino benzoic acid as corresponding to the 3-methylamino benzoic acid experimental mass spectrum and 4- methylamino benzoic acid as correctly corresponding to the 4-methylamino benzoic acid experimental mass spectrum.
[0179] Example 4
[0180] Example 4 sets out exemplary ways of calculating scores for assigning candidate isomers to the correct mass spectrum similarly to Examples 1 to 3 except that there are three candidate isomers in this example. In this example, the candidate isomers are 3,5-dimetoxyaniline, 2,4- dimetoxyaniline and 2, 6-dimetoxyaniline. A mass spectrum for 3,5-dimetoxyaniline, a mass spectrum for 2,4-dimetoxyaniline and a mass spectrum for 2, 6-dimetoxyaniline are each experimentally obtained using a Thermo Fisher GC / MS Orbitrap instrument. The candidate isomers are assigned to the correct experimental mass spectrum based on generating a score indicating the extent to which the candidate isomer correlates to the experimental mass spectrum. The score is calculated based on identifying m / z peaks in the associated mass spectral data that correspond to the m / z of the unique fragment(s) of the candidate isomer. Table 1 Candidate Isomers Structural Formula Data
[0181] The total number of fragments can be determined by performing fragmentation of the compound in-silico (i.e. by computer modelling or simulation). Performing such fragmentation in-silico based on the structural formula of the compound is well known in the art and may be performed using, for example, Mass Frontier. The fragments generated in-silico for that isomer that are not generated fort the other isomer are the unique fragments. In this case, 3,5-dimetoxyaniline has 29 unique fragments (fragments generated for 3,5-dimetoxyaniline but not generated for 2,4- dimetoxyaniline or 2,6-dimetoxyaniline) and 2,4-dimetoxyaniline has 6 unique fragments (fragments generated for 2,4-dimetoxyaniline but not generated for 3,5-dimetoxyaniline or 2,6- dimetoxyaniline) and 2,6-dimetoxyaniline has 15 unique fragments (fragments generated for 2,6- dimetoxyaniline but not generated for 2,4-dimetoxyaniline or 3,5-dimetoxyaniline)
[0182] Table 2: Total number of m / z peaks in the experimental mass spectra corresponding to the m / z ratios of unique fragments of the candidate isomers The 3,5-dimetoxyaniline experimental mass spectrum is analysed to determine whether there are any m / z peaks at m / z corresponding to the unique fragments of each isomer. In this case, the 3,5-dimetoxyaniline experimental mass spectrum has six m / z peaks at m / z ratios corresponding to the unique fragments of 3,5-dimetoxyaniline. The 3,5-dimetoxyaniline experimental mass spectrum also has one m / z peak at a m / z ratio corresponding to a unique fragment of 2,4- dimetoxyaniline. The 3,5-dimetoxyaniline experimental mass spectrum also has five m / z peaks at m / z ratios corresponding to unique fragments of 2,6-dimetoxyaniline. The 4-methylamino benzoic acid experimental mass spectrum has eight m / z peaks at m / z ratios corresponding to the unique fragments of 4-methylamino benzoic acid.
[0183] Ideally, there would be a non-zero value for number of m / z peaks corresponding to unique fragments of the first isomer in the corresponding experimental mass spectra and a zero value for number of m / z peaks corresponding to unique fragments of the first isomer in the noncorresponding mass spectrum. This is because each mass spectrum should only correspond to one of the isomers and so m / z peaks corresponding to unique fragments of that isomer should only be present in the mass spectrum of that isomer.
[0184] In this case, each of the experimental mass spectra include m / z peaks at m / z ratios corresponding to unique fragments for more than one isomer. This means that the in-silico fragmentation predicted a fragment for only one of the isomers and so identified this fragment as a unique fragment but such a fragment was generated for more than one of the isomers in practice (i.e. during experimental fragmentation). This is therefore a false negative prediction for that fragment. However, as discussed above, despite erroneous identification of a fragment as being unique for that isomer, the fragmentation model predicts correct fragments with a higher probability than incorrect ones, and the resulting score can consequently be used to correctly assign isomers to the corresponding mass spectrum, and so corresponding chromatogram peak.
[0185] The number of m / z peaks corresponding to the m / z of unique fragments of the isomers in the experimental mass spectra as set out in Table 2 are used to generate the scores set out in Tables 3 to 5.
[0186] Table 3: Absolute Intensity Ratio Score
[0187] The scores set out in Table 3 are the absolute intensity ratios of intensities of m / z peak(s) at m / z corresponding to the unique fragments of the candidate isomer to the sum of the absolute intensities of all m / z peaks in each experimental mass spectrum.
[0188] The 3,5-dimetoxyaniline experimental mass spectrum has a higher absolute intensity ratio for 3,5-dimetoxyaniline than the experimental mass spectrum for 2,4-dimetoxyaniline or the experimental mass spectrum for 2,6-dimetoxyaniline. Therefore, these absolute ratio scores can be used to correctly identify 3,5-dimetoxyaniline as corresponding to the 3,5-dimetoxyaniline experimental mass spectrum. The 3,5-dimetoxyaniline experimental mass spectrum, the 2,4- dimetoxyaniline experimental mass spectrum and the 2,6-dimetoxyaniline experimental mass spectrum have the same absolute intensity ratios for 2,4-dimetoxyaniline and 2,6-dimetoxyaniline and so the absolute intensity ratio scores cannot be used to identify which experimental mass spectrum corresponds to 2,4-dimetoxyaniline or 2,6-dimetoxyaniline.
[0189] Table 4: Relative Intensity Ratio Score
[0190] The scores set out in Table 4 are the relative (normalised) ratio of intensities of m / z peak(s) at m / z ratios corresponding to the unique fragments of the candidate isomer to intensities of all m / z peaks in each experimental mass spectrum.
[0191] The 3,5-dimetoxyaniline experimental mass spectrum has a higher relative intensity ratio for candidate isomer 3,5-dimetoxyaniline than the 2,4-dimetoxyaniline experimental mass spectrum or the 2, 6 3,5-dimetoxyaniline experimental mass spectrum. The 3,5-dimetoxyaniline experimental mass spectrum has the same relative intensity ratio for candidate isomer 2, 4-dimetoxyaniline as the 2,4-dimetoxyaniline experimental mass spectrum and the 2, 6-dimetoxyaniline experimental mass spectrum.
[0192] The 3,5-dimetoxyaniline experimental mass spectrum has the same relative intensity ratio for candidate isomer 2, 6-dimetoxyaniline as the 2,4-dimetoxyaniline experimental mass spectrum and the 2, 6-dimetoxyaniline experimental mass spectrum.
[0193] Therefore, these relative intensity ratios correctly identify 3,5-dimetoxyaniline as corresponding to the 3,5-dimetoxyaniline experimental mass spectrum but do not identify which experimental mass spectrum 2,4-dimetoxyaniline or 2, 6-dimetoxyaniline correspond to.
[0194] Table 5: Absolute Peak Count Ratio Score
[0195] The scores set out in Table 5 are the absolute peak count ratios of the number of m / z peak(s) corresponding to m / z ratios of the unique fragments of the candidate isomer to the number of all m / z peaks in each experimental mass spectrum.
[0196] The 3,5-dimetoxyaniline experimental mass spectrum has a higher absolute peak count ratio for candidate isomer 3,5-dimetoxyaniline than the 2,4-dimetoxyaniline experimental mass spectrum or the 2, 6-dimetoxyaniline experimental mass spectrum.
[0197] The 2, 6-dimetoxyaniline experimental mass spectrum has a higher absolute peak count ratio for candidate isomer 2, 6-dimetoxyaniline than the 3,5-dimetoxyaniline experimental mass spectrum or the 2,4-dimetoxyaniline experimental mass spectrum.
[0198] Therefore, these absolute peak count ratios correctly identify 3,5-dimetoxyaniline as corresponding to the 3,5-dimetoxyaniline experimental mass spectrum and 2, 6-dimetoxyaniline as correctly corresponding to the 2, 6-dimetoxyaniline experimental mass spectrum. By the process of elimination 2,4-dimetoxyaniline is determined as correctly corresponding to the 2,4- dimetoxyaniline experimental mass spectrum.
[0199] In summary, in Example, 4, each of the three candidate isomers have been accurately assigned to the correct corresponding experimental mass spectrum. None of the assignments made were incorrect.
[0200] In summary, in all of these examples employing two isomers, at least one isomer was correctly assigned to its corresponding experimental mass spectrum. Subsequently, the second isomer can also be assigned to the correct spectrum, at least by the process of elimination. In the example employing three isomers, at least two isomers were correctly assigned to their corresponding experimental mass spectrum. Subsequently, the third isomer can also be assigned to the correct spectrum, at least by the process of elimination. In any of the examples, there is no incorrect assignment of a candidate compound to an experimental mass spectrum.
[0201] As exemplified by Example 4, in the event that there are more than two candidate isomers to assign to chromatogram peaks based on mass spectral data, the scores are calculated in the same way and the candidate isomer is assigned to the mass spectral data for which the score is the highest. The scores can be represented in a table where all isomers are set out in the rows and the mass spectral data set out in the columns. The highest score value in one row assigns the candidate isomer to the most probable mass spectrum and so most probable chromatogram peak
[0202] It will be appreciated that these examples demonstrate that this scoring method can be applied to assign candidate isomers to the mass spectral data associated with a chromatogram peak.
[0203] Step 103 - Option (i)
[0204] Step 103, option (i) is best described in the context of the method steps set out in Figure 2.
[0205] Steps 201 and 202 of Figure 3 are the same as steps 101 and 102 of Figure 2, which are described above.
[0206] After the step of obtaining associated mass spectral data for each peak in the chromatogram, the method further comprises the steps of: obtaining a list of candidate chemical structures for the compounds based on the associated mass spectrum for each peak and allocating each of the candidate chemical structures to one or more of the chromatogram peaks based on the associated mass spectrum of the chromatogram peak (step 203a). The list of candidate chemical structures for the compounds are obtained based on the associated mass spectral data for each peak. A database may be searched for candidate chemical structures having mass spectral data corresponding to the mass spectral data associated with each chromatogram peak. A compound in the database may be identified as a candidate structural formula when mass spectral data of the compound in the database corresponds with the associated mass spectral data of one of the chromatogram peaks.
[0207] The associated mass spectral data for each peak may comprise one or more MSnmass spectra where n > 1 . The associated mass spectral data may be one of a plurality of mass spectra associated with the chromatogram peak. For example, each peak may have an associated MS1spectrum and MS2spectrum. Each peak may have a plurality of multi-stage MSnassociated mass spectra that may form a mass spectral tree. The mass spectral data may be a list of identified m / z peaks in the one or more MSnmass spectra where n > 1 .
[0208] The database may be a mass spectral library that contains compounds defined according to their structural formula together with their mass spectral data.
[0209] The mass spectral data for each compound stored in the database may similarly comprise one or more MSnmass spectra where n > 1 . The mass spectral data may be a list of one or more m / z peaks in the one or more MSnmass spectra where n > 1 . The mass spectral data may include m / z value(s) associated with the compound, for example m / z peak(s) of ion(s) of the compound. By way of example, the mass spectral data stored in the database for a compound may be the spectral fingerprint of the compound. By way of example, the mass spectral data for each compound stored in the databased may be represented as one or more mass spectral tree(s).For example, each tree may be comprise different spectral nodes. The spectral nodes can be a different level MSnionization or may be the same level of MSnbut may have been ionized by different cation (H,Na,K,) Each spectral node may comprise one or more spectra. Each spectra may be a list of one or more m / z peaks. Other ways of representing the mass spectral data are contemplated.
[0210] The search of the database may be based on one or more of the mass spectra or based on m / z peaks within the mass spectra above a certain threshold or based on the mass spectral tree associated with the chromatogram peak.
[0211] For example, the database may be searched for candidate chemical structures having mass spectra similar to the mass spectra corresponding to each chromatogram peak. By way of further example, the database may be searched for candidate chemical structures having a mass spectral tree similar to the mass spectral tree corresponding to each chromatogram peak.
[0212] By way of further example, the database may be searched for candidate chemical structures having mass spectral data comprising peaks at m / z ratios corresponding to (the same as or within a tolerance range of) m / z ratios of peaks of the associated mass spectral data of each chromatogram peak.
[0213] By way of further example, the database may be searched for candidate chemical structures having m / z ratio(s) correspond to (the same as or within a tolerance range of) m / z ratios of peaks of the associated mass spectral data of each chromatogram peak.
[0214] Searching the database for candidate chemical structures having mass spectral data corresponding to the mass spectral data associated with each chromatogram peak may comprise identifying m / z ratios of one or more m / z peak(s) in each associated mass spectrum; and identifying a compound in the database as a candidate chemical structure when m / z value(s) associated with the compound, for example m / z peak(s) of ion(s) of the compound in the database corresponds with one or more of the identified m / z ratio(s).. Identifying one or more m / z peaks in each associated mass spectrum may involve identifying the m / z ratios of all m / z peaks within each associated mass spectrum above a threshold. By “corresponds” in this context, this refers to the m / z ratio of the compound in the database being the same as or within a tolerance range of the m / z ratio(s) of m / z peaks of the associated mass spectrum of the chromatogram peaks.
[0215] Examples of databases are mzCloud or NIST. The chemical structure representations in the mass spectral library may be graph-based and / or descriptor-based.
[0216] In one example, each of the mass spectral data of the chromatogram peaks may be input into the database to obtain the list of candidate compounds for each chromatogram peak. A separate search may be performed for each chromatogram peak.
[0217] The candidate chemical structure(s) found in the database for each chromatogram peak are allocated to that chromatogram peak. For each chromatogram peak, a plurality of candidate chemical structures may be identified (i.e. the search results in a plurality of hits) and allocated to that chromatogram peak. Alternatively, instead of performing a separate search for each chromatogram peak, all of the mass spectral data of all of the chromatogram peaks may be input into the database to obtain the list of candidate compounds for all of the chromatogram peaks by way of a single search.
[0218] Candidate chemical structures that are also isomers (compounds having the same molecular formula but a different structural formula) are referred to as candidate isomers herein. As isomers have the same mass and so the same molecular ion m / z ratio (the m / z peak having the highest mass), it is difficult to rely on such a database search to distinguish which isomer corresponds to which chromatogram peak. Consequently, typically isomers are allocated to multiple chromatogram peaks.
[0219] In the general description of the method set out in Figure 2, the list of candidate chemical structures comprises a first plurality of candidate chemical structures and fragments are predicted for each candidate chemical structure of the first plurality of candidate chemical structures.
[0220] In step 203b, chromatogram peaks having candidate chemical structures allocated thereto that are candidate isomers with a first molecular formula are grouped together into a first group. As discussed above, candidate chemical structures having a first molecular formula may be referred to herein as candidate chemical structures of the first plurality of candidate chemical structures. The chromatogram peaks within the first group may optionally have one or more candidate chemical structure(s) allocated thereto that are not candidate isomer(s) in addition to the candidate chemical isomers with the first molecular formula. Alternatively, the chromatogram peaks within the first group may only have candidate chemical isomers with the first molecular formula allocated thereto.
[0221] Step 203b may comprise grouping the chromatogram peaks into a plurality of groups where each group of chromatogram peaks have candidate chemical structures allocated thereto that have the same molecular formula.
[0222] For example, step 203b may further comprise grouping chromatogram peaks having candidate chemical structures allocated thereto that are candidate chemical structures with a second molecular formula into a second group where the second molecular formula is different from the first molecular formula. Candidate chemical structures having a second molecular formula may be referred to herein as the second plurality of candidate chemical structures. The second plurality of candidate chemical structures may consist of candidate isomers having the second molecular formula. The chromatogram peaks within the second group may optionally have one or more candidate chemical structure(s) allocated thereto that are not candidate isomer(s) in addition to the candidate chemical isomers with the second molecular formula. Alternatively, the chromatogram peaks within the first group may only have candidate chemical isomers with the second molecular formula allocated thereto.
[0223] In an embodiment employing first and second groups, the list of candidate chemical structures in step 203a may comprise the first plurality of candidate chemical structures and the second plurality of candidate chemical structures. The list of candidate chemical structures in step 203a may also comprise candidate chemical structures that do not have the first or second molecular formula.
[0224] In the exemplary embodiment shown in Figure 3, in step 203c, fragments of at least each of the candidate isomers having the first molecular formula are then predicted. The fragments are predicted by performing fragmentation of the candidate chemical structure in-silico (i.e. by computer modelling or simulation). Performing such fragmentation in-silico based on the structural formula of the compound is well known in the art and may be performed using, for example, Mass Frontier.
[0225] Unique fragments for each candidate isomer having the first molecular formula and their m / z, are then identified (step 203c). As discussed above, a unique fragment is a fragment predicted for a candidate chemical structure within the first plurality of candidate chemical structures that is not predicted for the other candidate chemical structures within the first plurality of candidate chemical structures. In the method of Figure 3, the first plurality of candidate chemical structures optionally consists of candidate isomers having a first molecular formula. Therefore, a unique fragment is a fragment of a candidate isomer having the first molecular formula that is not predicted for the other candidate isomers having the first molecular formula. A unique fragment of a candidate isomer having the first molecular formula is a fragment that has a m / z ratio that is different from fragments generated for the other candidate isomers having the first molecular formula.
[0226] In step 304, a candidate isomer having the first molecular formula is assigned to a chromatogram peak within the first group (specifically the most probable chromatogram peak within the first group) based on identifying one or more m / z peak(s) corresponding to the m / z of the unique fragment in the mass spectral data associated with the chromatogram peak.
[0227] Multiple candidate isomers within the first plurality of candidate chemical structures may be assigned to multiple chromatogram peaks within the first group such that each chromatogram peak within the first group only has one candidate isomer assigned thereto. The step of, for one or more of the candidate isomers having the first molecular formula, assigning the candidate isomer to a chromatogram peak within the first group (step 204) may comprise determining a score for each chromatogram peak within the first group indicating the extent to which the candidate isomer correlates to the associated mass spectral data of the chromatogram peak within the first group and assigning the candidate isomer to the chromatogram peak within the first group having the score indicating the greatest correlation. The score may be calculated based on identifying one or more m / z peak(s) in the associated mass spectral data that correspond to the m / z ratio of the unique fragment(s) of the candidate isomer. The associated mass spectral data comprises at least one associated MSnmass spectrum. The score may be calculated based on identifying one or more m / z peak(s) in the at least one associated mass spectrum at m / z corresponding to the m / z of the unique fragment(s) of the candidate isomer.
[0228] Calculating such scores is described in respect of step 104 of Figure 1 and demonstrated by the above Examples 1 , 2, 3 and 4 which equally apply to step 204 of the method of this embodiment.
[0229] In an exemplary embodiment, the list of candidate chemical structures obtained based on the associated mass spectral data of the chromatogram peaks may comprise a second plurality of candidate chemical structures that consist of candidate isomers having the second molecular formula in addition to the first plurality of candidate chemical structures consisting of candidate isomers having the first molecular formula. The first molecular formula being different from the second molecular formula. Some chromatogram peaks may have multiple candidate chemical structures of the first plurality of candidate chemical structures allocated thereto based on the associated mass spectral data. Some chromatogram peaks may have multiple candidate chemical structures of the second plurality of candidate chemical structures allocated thereto based on the associated mass spectral data.
[0230] As discussed above, the step of grouping together the chromatogram peaks having allocated candidate chemical structures that are candidate isomers with same molecular formula (step 203b) may comprise grouping together chromatogram peaks having allocated candidate isomers with the first molecular formula into a first group and grouping together isomer chromatogram peaks having allocated candidate isomers with the second molecular formula into a second group.
[0231] Steps 203c and 204 would be performed for the candidate isomers within the first plurality of candidate chemical structures. The method would further comprise performing similar steps in respect of the candidate isomers within the second plurality of candidate chemical structures. In other words, the method would further comprise predicting fragments of each candidate isomer within the second plurality of candidate chemical structures and identifying a m / z ratio of at least one unique fragment for each candidate chemical structure of the second plurality of candidate chemical structures (a fragment that is unique compared to the fragments of other candidate isomers within the second plurality of candidate chemical structures). The method would then comprise, for one or more of the candidate isomer(s) of the second plurality of candidate chemical structures, assigning the candidate isomer to a chromatogram peak within the second group based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z of the at least one unique fragment(s) of the candidate isomer.
[0232] Such an embodiment may be applicable to a sample comprising a mixture of a plurality of compounds where the mixture of the plurality of compounds comprises (optionally consists of) one set of isomers having the first molecular formula and one set of isomers having the second molecular formula. The list of candidate chemical structures may similarly consist of one set of isomers (the first plurality of candidate chemical structures having the same (the first) molecular formula but different structural formulas).
[0233] The order of the method steps shown in Figure 3 may be varied provided the context allows and some of the method steps may be performed simultaneously. The steps of obtaining the chromatogram and associated mass spectral data needs to be performed the steps of obtaining the list of candidate chemical structures. The step of obtaining the list of candidate chemical structures needs to be performed before the step of allocating the candidate chemical structures and grouping the chromatogram peaks according to the allocated candidate chemical structures. The steps of allocating the candidate chemical structures, grouping the chromatogram peaks according to the allocated candidate chemical structures, predicting fragments of the candidate isomers and identifying a m / z of at least one unique fragment(s) of candidate isomers need to be performed before the step of assigning the candidate isomer(s) to chromatogram peak(s). Otherwise, the method steps may be performed in any order or sequentially.
[0234] One or more of the method steps performed in respect of the first plurality of candidate chemical structures may be before simultaneously or sequentially to the method steps performed in respect of the second plurality of candidate chemical structures.
[0235] An exemplary implementation of method steps 201 , 202, 203a, 203b and 203c in respect of such an embodiment where the sample contains at least two compounds that are isomers having a first molecular formula and at least two compounds that are isomers having a second molecular formula is schematically depicted in Figure 4. In the exemplary embodiment of the method schematically shown in Figure 4 a chromatogram of the sample is obtained in step 201. The chromatogram is sketched schematically as having four peaks, A,B,C, and D.
[0236] In step 202 of Figure 4, mass spectral data associated with each chromatogram peak is obtained. The mass spectral data sets out a characteristic m / z peak for each chromatogram peak. Such a characteristic peak is the molecular ion peak (the m / z peak having the highest m / z ratio). The characteristic m / z mass peak for chromatogram peak A is at a m / z of 271 .060. The characteristic m / z peak for chromatogram peak B is at a m / z of 371 .066. The characteristic m / z peak for chromatogram peak C is at a m / z of 271 .060. The characteristic m / z peak for chromatogram peak D is at a m / z of 371 .066.
[0237] In step 203a of Figure 4, a search of a mass spectral database was performed using the associated mass spectral data of each chromatogram peak to obtain a list of candidate chemical structures. In this example, the database employed was NIST, and the search performed was for compounds having mass spectral data containing m / z peak(s) at the same value(s) as the characteristic m / z peak(s) of the mass spectral data of each chromatogram peak. The candidate chemical structures identified were: apigenin and genistein for chromatogram peak A;
[0238] Isohamnetin , Tamarixetin and Isorhamnetin for chromatogram peak B; apigenin and genistein for chromatogram peak C; and Isohamnetin, Tamarixetin and Isorhamnetin for chromatogram peak D.. As these candidate chemical structures are also isomers, they are referred to as candidate isomers.
[0239] In step 203b of Figure 4, peaks A and C are grouped together, since they have allocated candidate chemical structures that are candidate isomers having the same molecular formula. Peaks B and D are grouped together, since they also have allocated candidate chemical structures that are candidate isomers having the same molecular formula.
[0240] In step 203c, fragments of each candidate chemical structure are predicted in-silico (by computer modelling or simulation), for example, using Mass Frontier. Unique fragments of each candidate chemical structure within the first plurality of candidate chemical structures and unique fragments of each candidate chemical structure within the second plurality of candidate chemical structures are then identified. In other words, unique fragments of the candidate chemical structures within each group are identified. Unique fragments are those fragments that are predicted for a candidate chemical structure within the but not predicted for the other candidate chemical structure(s) within the same group. Those unique fragments are represented in step 203c of Figure 5 according to their 2D structural formula. The method would then comprise step 204 (not shown in Figure 5) where each candidate chemical structure is assigned to the most probable chromatogram peak within the group.
[0241] For the candidate chemical structure, one or more score(s) is determined for each isomer chromatogram peak within the group indicating the extent to which the candidate chemical structure within the group correlates to the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak within the group having the score indicating the greatest correlation. The score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak that correspond to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure. The scoring methods described above in respect of step 104 and demonstrated by Examples 1 , 2, 3 and 4 above equally apply to step 204 of Figure 5. The scoring is performed for each candidate chemical structure so as to assign each candidate chemical structure to the most probable chromatogram peak. Consequently, each chromatogram peak has a single candidate chemical structure assigned thereto. As shown in Examples 1 , 2, 3 and 4, the scores may be arranged within a table. In this case, the columns would correspond to the mass spectral data, or mass spectrum of each chromatogram peak within the group and the rows would correspond to the candidate chemical structures within the group.
[0242] Step 103 option (ii)
[0243] Step 103, option (ii) is best described in the context of the method steps set out in Figure 5.
[0244] Steps 301 and 302 of Figure 5 are the same as those of Figure 1 except that the mixture of compounds of the sample comprise derivatives of a known precursor. In other words, some of the compounds within the sample are derivatives of a known precursor. Optionally, the sample consists of (comprises only) derivatives of a known precursor. Optionally, the sample consists of (comprises only) derivatives of a known precursor and the known precursor.
[0245] Derivatives of a precursor are compounds resulting from derivatisation of the precursor. Derivatization is the chemical reaction between the precursor and a derivatisation agent to form the derivatives. Functional groups of the known precursor react with the derivatization agent to form derivatized groups. The precursor may comprises a plurality of functional groups. The functional groups may be referred to as active functional groups. Some examples of functional groups include: alcohols, alkenes, alkynes, amines, carboxylic acids, aldehydes, ketones, esters, active hydrogens and ethers. If the precursor has a plurality of functional groups, then the derivatization reaction can result in a mixture of derivatives with different degrees of derivatization and various positions of the derivatized group within the structure. Some of those derivatives may have the same degree of derivatization but different positions of the derivatized group within the structure thereby resulting in isomers (compounds having the same molecular formula but different structural formula).
[0246] Derivatisation and the use of derivatisation agents is well known in the art. In known chromatography techniques, derivatization is often performed to input to the chromatography column improve, for example, the chromatographic separation and / or thermal stability of the compound.
[0247] Derivatization can involve, for example, silylation, aklylation or acylation reactions. During derivitization, the derivitization agent reacts with an active functional group in the compound, for example active hydrogens in the case of silylation. Derivatisation agents include, for example, N,O-bis(trimethylsilyl)tri fluoroacetamide,
[0248] 1 ,3-bis(chloromethyl)-1 ,1 ,3,3-tetramethyldisilazane, dimethylchlorosilane, N- heptafluorobutyrylimidazole.
[0249] The method may optionally involve performing derivatisation of the known compound to obtain the sample and then performing chromatography in respect of the sample to obtain the chromatogram. As discussed above, the chromatography may be gas or liquid chromatography.
[0250] Step 303a comprises obtaining a list of candidate chemical structures for the compounds, wherein the list of candidate chemical structures comprise predicted derivatives of the known precursor, wherein at least two of the predicted derivatives are isomers. The list of candidate chemical structures may consist of (i.e. comprise only) the predicted derivatives of the known precursor, and optionally, also the known precursor.
[0251] The list of candidate chemical structures comprises the first plurality of candidate chemical structures, as described above. As discussed above, the first plurality of candidate chemical structures are candidate chemical structures for which fragments are predicted.
[0252] In this exemplary embodiment of Figure 5, the first plurality of candidate chemical structures comprise, and optionally may consist of (i.e. comprise only), the predicted derivatives of the known precursor. The first plurality of candidate chemical structures may consist of (i.e. comprise only) the known precursor (i.e. the known precursor in its underivatized form) and predicted derivatives of the known precursor. In embodiments where the sample has been produced using an excess of derivatisation agent, it may be expected that the precursor is fully derivatized and so the sample and the candidate chemical structures may not include the known precursor. However, in embodiments, where an excess of derivatisation agent has not been used, then it may be expected that the precursor is not fully derivatised and so the sample and the list of candidate chemical structures and / or the first plurality of candidate chemical structures may include the known precursor.
[0253] The structural formulas of the predicted derivatives may be determined by simulating the derivatization reaction between the known precursor and derivatization agent using computer modelling (i.e. in-silico). It is possible to use a software package to simulate chemical reactions in-silico based on input of the 2D structural formula of the reagents to be reacted (in this case the known precursor and the derivatisation agent). An exemplary software package is RDKit. The chemical reactions may be constructed using, for example, a SMARTS-based naming format, similar to Daylight Reaction’s SMILES or SMIRKS as is known in the art. The predicted derivatives may be stored in a database in association with the known precursor and derivatization agent.
[0254] Alternatively, the structural formula of each of the predicted derivatives may have been previously determined by generating derivatives of the known compound by experimentation. The structural formula of each of the predicted derivatives of the known compound may be prestored in a database in association with the known precursor and derivatization agent and obtained from that database.
[0255] The predicted derivatives that are also isomers may be referred to as the candidate isomers. Optionally, all of the predicted derivatives may be isomers (i.e. compounds having the same molecular formula but different structural formula). In particular, the first plurality of candidate chemical structures may consist of predicted derivatives that are candidate isomers.
[0256] As discussed above, the first plurality of candidate chemical structures consist of (i.e. comprise only) predicted derivatives of the known precursor and so the rest of the method will be discussed in the context of the predicted derivatives. In step 303b, fragments of the predicted derivatives of the known precursor are predicted by performing fragmentation of the predicted derivatives in-silico (i.e. by computer modelling or simulation). Performing such fragmentation in- silico based on the structural formula of the compound is well known in the art and may be performed using, for example, Mass Frontier. The fragments generated in-silico for one predicted derivative that are not generated for the other predicted derivatives are the unique fragments. The m / z ratios of the unique fragment(s) for each predicted derivative are then obtained, for example, using mass frontier. Each predicted derivative may have a plurality of unique fragments. The fragments predicted in-silico may not all be produced in reality as explained in Examples 1 , 2,3 and 4.
[0257] Step 304 is then performed that comprises, for one or more of the predicted derivatives, assigning the predicted derivative to a chromatogram peak (the most probable chromatogram peak) based on identifying one or more m / z peak(s) in the associated mass spectral data at m / z ratios corresponding to the m / z of the at least one unique fragment(s) of the predicted derivative. Each chromatogram peak may have a single predicted derivative assigned thereto.
[0258] As discussed above, the step of assigning a candidate chemical structure (in this example a predicted derivative) to a chromatogram peak may comprise determining a score indicating the extent to which the predicted derivative correlates / corresponds to the associated mass spectral data of each chromatogram peak and assigning the predicted derivative to chromatogram peak having the score indicating the greatest correlation. In other words, the score indicates the likelihood of the predicted derivative corresponding to the chromatogram peak and the predicted derivative is assigned to the chromatogram peak having the score indicative of the highest likelihood. The scoring methods described above equally apply to this embodiment.
[0259] The order of the method steps shown in Figure 4 may be varied provided the context allows and some of the method steps may be performed simultaneously. Step 303a and 303b may be performed simultaneously with or before or after Steps 301 and 302.
[0260] Figure 6 depicts an exemplary implementation of steps 303a and 303b of the method of Figure 5 where the known precursor is cimifugin and the predicted derivatives are isomers, referred to herein as candidate isomers. The structural formula of the predicted derivatives (first and second derivatives) are shown in Figure 6. The first predicted derivative having an m / z of 378.1499. The second predicted derivative having an m / z of 378.1499. Fragments of each predicted derivative are predicted by performing fragmentation of each predicted derivative in-silico. The unique fragments (fragments predicted for one predicted derivative but not for the other predicted derivatives) are then identified in Step 303b and their 2D structural formula is shown in Figure 6. The unique fragments for the first derivative having m / z values of 117.0730 and 323.1309.
[0261] Figure 7(a) includes a mass spectrum of a sample containing a mixture of the first and second derivatised forms of cimifugin, the mass spectrum including m / z peaks corresponding to each of the unique fragments. Figure 7(a) includes a mass spectrum for the first derivatised form of cimifugin, the mass spectrum including m / z peaks corresponding to the unique fragments of the first derivatised form (at m / z values of 117 and 323) but not including m / z peaks corresponding to the unique fragments of the second derivatised form (at m / z values of 331 and 319). The m / z peaks at 117 are more evident in the zoomed-in version of the mass spectra provided in Figure 7(b). The m / z peaks at 323 are more evident in the zoomed-in version of the mass spectra provided in Figure 7(c). Figure 7(a) includes a mass spectrum for the second derivatised form, the mass spectrum including m / z peaks corresponding to the unique fragments of the second derivatised form (at m / z values of 331 and 319) but not m / z peaks corresponding unique fragments of the first derivatised form (at m / z values of 117and 323). The m / z peaks at 331 and 319 are more evident in the zoomed-in version of the mass spectra provided in Figure 7(c). In each of Figures 7(a) to (c), the middle mass spectrum is of a sample containing a mixture of first and second derivatives of cimifugin, the bottom mass spectrum is of the first derivatised form and the top mass spectrum is of the second derivatised form. In Figures 7(b) and 7(c) the m / z peaks corresponding to unique fragments are identified with boxes therearound.
[0262] The methods described herein may be implemented as a computer program or programmable or programmed logic configured to perform the method when operated by a processor. In other words, the methods described herein may be implemented as one or more corresponding modules as hardware and / or software. For example, the methods may be implemented as one or more software components for execution by a processor of the system. Alternatively, the above- mentioned functionality may be implemented as hardware, such as on one or more field- programmable-gate-arrays (FPGAs), and / or one or more application-specific-integrated-circuits (ASICs), and / or one or more digital-signal-processors (DSPs), and / or other hardware arrangements.
[0263] It will be appreciated that, insofar as embodiments of the disclosure are implemented by a computer program, then a storage medium and a transmission medium carrying the computer program form aspects of the disclosure. The computer program may have one or more program instructions, or program code, that, when executed by a processor, causes an embodiment of the disclosure to be carried out. The term “program”, as used herein, may be a sequence of instructions designed for execution on a processor, and may include a subroutine, a function, a procedure, a module, an object method, an object implementation, an executable application, an applet, a servlet, source code, object code, a shared library, a dynamic linked library, and / or other sequences of instructions designed for execution on a computer system. The storage medium may be a magnetic disc (such as a hard drive or a floppy disc), an optical disc (such as a CD-ROM, a DVD-ROM or a BluRay disc), or a memory (such as a ROM, a RAM, EEPROM, EPROM, Flash memory or a portable / removable memory device), etc. The transmission medium may be a communications signal, a data broadcast, a communications link between two or more computers, etc. The computer program comprising instructions that, when executed by a processor, cause the processor to perform the methods described herein may form part of a data analysis system. The data analysis system may further comprise a memory storing data, such as the chromatogram to be analysed. The data analysis system may be part of an apparatus comprising a chromatography instrument and mass spectrometer. The chromatography instrument may be a gas chromatography instrument or a liquid chromatography instrument. Such chromatography instruments are well known in the art. Typically, chromatography instruments comprise a chromatography column configured to separate the sample into its components, an injector configured to inject the sample into the chromatography column and a detector configured to detect the separated components. The mass spectrometer may be any type of mass spectrometer known in the art. The data analysis system may be configured to receive the chromatogram from the chromatography instrument.
[0264] As used herein, including in the claims, unless the context indicates otherwise, singular forms of the terms herein are to be construed as including the plural form and vice versa. For instance, unless the context indicates otherwise, a singular reference herein including in the claims, such as "a" or "an" (such as an analogue to digital convertor) means "one or more" (for instance, one or more analogue to digital convertor). Throughout the description and claims of this disclosure, the words "comprise", "including", "having" and "contain" and variations of the words, for example "comprising" and "comprises" or similar, mean "including but not limited to", and are not intended to (and do not) exclude other components.
[0265] The use of any and all examples, or exemplary language ("for instance", "such as", "for example" and like language) provided herein, is intended merely to better illustrate the disclosure and does not indicate a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non- claimed element as essential to the practice of the disclosure.
[0266] Any steps described in this specification may be performed simultaneously unless stated or the context requires otherwise.
[0267] All of the aspects and / or features disclosed in this specification may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. As described herein, there may be particular combinations of aspects that are of further benefit, for example in respect of the jet disruptor in conjunction with a sealing valve and / or an offset of the axial bore and downstream aperture. In particular, the preferred features of the disclosure are applicable to all aspects of the disclosure and may be used in any combination. Likewise, features described in non-essential combinations may be used separately (not in combination).
Claims
CLAIMS1 . A method of analysing a chromatogram of a sample, the method comprising: obtaining a chromatogram of the sample, the sample comprising a mixture of a plurality of compounds, wherein at least two of the compounds are isomers, the chromatogram comprising a plurality of chromatogram peaks; obtaining associated mass spectral data for each chromatogram peak of the plurality of chromatogram peaks; predicting fragments of each candidate chemical structure of a first plurality of candidate chemical structures and identifying a m / z ratio of at least one unique fragment of each candidate chemical structure of the first plurality of candidate chemical structures, wherein at least two of the candidate chemical structures of the first plurality of candidate chemical structures are candidate isomers; and for one or more of the candidate chemical structure(s) of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak of the plurality of chromatogram peaks based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio(s) of the at least one unique fragment(s) of the candidate chemical structure, wherein at least one of the assigned candidate chemical structure(s) is one of the candidate isomers.
2. The method of claim 1 , wherein one or more of the candidate chemical structures of the first plurality of candidate chemical structures has a plurality of unique fragments.
3. The method of claim 1 , wherein all of the candidate chemical structures of the first plurality of candidate chemical structures are candidate isomers.
4. The method of any preceding claim, wherein each of the associated mass spectral data comprises an associated MSnmass spectrum where n >1 , optionally wherein each of the associated mass spectral data comprises a plurality of associated MSnmass spectrum where n >1 , further optionally wherein each of the associated mass spectral data comprises multi-stage MSnmass spectra forming a spectral tree.
5. The method of any preceding claim, wherein the step of assigning, for one or more of the candidate chemical structure(s) of the first plurality of candidate chemical structures, the candidate chemical structure to a chromatogram peak, is performed for each of the first plurality of candidate chemical structures.
6. The method of any preceding claim, further comprising obtaining a list of candidate chemical structures comprising the first plurality of candidate chemical structures.
7. The method of claim 6, wherein the list of candidate chemical structures are obtained using a database, optionally wherein the database is a mass spectral library.
8. The method of claim 6 or claim 7, wherein the list of candidate chemical structures are obtained based on (i) the associated mass spectral data of each chromatogram peak and / or based on (ii) knowledge of a precursor to one or more compound(s) of the plurality of compounds or knowledge of a molecular formula of one or more compound(s) of the plurality of compounds.
9. The method of claim 7, wherein the list of candidate chemical structures are obtained based on the associated mass spectral data of each chromatogram peak, wherein obtaining the list of identified candidate chemical structures comprises searching the database for candidate chemical structures having mass spectral data corresponding to the mass spectral data associated with each chromatogram peak.
10. The method of claim 9, wherein the associated mass spectral data for each chromatogram peak comprises an associated mass spectrum, wherein searching the database for candidate chemical structures having mass spectral data corresponding to the mass spectral data associated with each chromatogram peak comprises: identifying one or more m / z peak(s) in the associated mass spectrum; and identifying a compound in the database as a candidate chemical structure when one or more m / z ratio(s) associated with the compound in the database corresponds with m / z ratio(s) of one or more of the identified m / z peak(s) in the associated mass spectrum.11 . The method of any one of claims 6 to 10, wherein the list of candidate chemical structures are obtained based on the associated mass spectral data of each chromatogram peak, wherein the method further comprises, before at least the step of assigning the candidate chemical structure to a chromatogram peak, the step of: allocating candidate chemical structures, including the first plurality of candidate chemical structures, from the list of candidate chemical structures, to one or more of thechromatogram peak(s) based on the associated mass spectral data of the chromatogram peaks, wherein the candidate isomers are allocated to multiple chromatogram peaks.
12. The method of claim 11 , wherein the first plurality of candidate chemical structures consists of candidate isomers having a first molecular formula, wherein the method further comprises, before at least the step of assigning the candidate chemical structure to a chromatogram peak, grouping together chromatogram peaks of the plurality of chromatogram peaks having one or more candidate chemical structure(s) of the first plurality of candidate chemical structures allocated thereto into a first group; wherein the step of, for one or more candidate chemical structure(s) of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak, comprises assigning the candidate chemical structure to a chromatogram peak within the first group based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratios corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure.
13. The method of claim 12, wherein the list of candidate chemical structures comprises a second plurality of candidate chemical structures, wherein the second plurality of candidate chemical structures consists of candidate isomers having a second molecular formula, wherein the second molecular formula is different from the first molecular formula, wherein the method further comprises grouping together chromatogram peaks having one or more candidate chemical structure(s) of the second plurality candidate chemical structures allocated thereto into a second group, wherein the method further comprises predicting fragments of each candidate chemical structure of the second plurality of candidate chemical structures and identifying a m / z ratio of at least one unique fragment of each candidate chemical structure of the second plurality of candidate chemical structures, wherein at least two of the candidate chemical structures of the second plurality of candidate chemical structures are candidate isomers; and for one or more of the candidate chemical structure(s) of the second plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak of the plurality of chromatogram peaks within the second group based on identifying one or more m / z peak(s) in the associated mass spectral data of the chromatogram peak at m / z ratio(s) corresponding to the m / z ratio(s) of the at least oneunique fragment(s) of the candidate chemical structure, wherein at least one of the assigned candidate chemical structure(s) is one of the candidate isomers.
14. The method of any one of claims 6 to 8, wherein the list of candidate chemical structures are obtained based on knowledge of a precursor to one or more compound(s) of the plurality of compounds, wherein the first plurality of candidate compounds consist of predicted derivatives of the known precursor, wherein at least two of the predicted derivatives are candidate isomers, optionally wherein the predicted derivatives are determined by generating derivatives of the known precursor in-silico.
15. The method of claim 14, wherein the method further comprises performing derivatisation of the known precursor to obtain the sample and performing chromatography of the sample to obtain the chromatogram.
16. The method of any one of claims 1 to 11 , 14 or 15, wherein the step of, for one or more candidate chemical structures of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak of the plurality of chromatogram peaks, comprises determining a score for each chromatogram peak indicating the extent to which the candidate chemical structure correlates to the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak having the score indicating the greatest correlation, wherein the score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure.
17. The method of claim 12, wherein the step of, for one or more candidate chemical structures of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak within the first group, comprises determining a score for each chromatogram peak within the first group indicating the extent to which the candidate chemical structure correlates to the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak within the first group having the score indicating the greatest correlation, wherein the score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure.
18. The method of claim 13, wherein the step of, for one or more candidate chemical structures of the first plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak within the first group, comprises determining a score for each chromatogram peak within the first group indicating the extent to which the candidate chemical structure correlates to the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak within the first group having the score indicating the greatest correlation, wherein the score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure, wherein the step of, for one or more candidate chemical structures of the second plurality of candidate chemical structures, assigning the candidate chemical structure to a chromatogram peak within the second group, comprises determining a score for each chromatogram peak within the second group indicating the extent to which the candidate chemical structure correlates to the associated mass spectral data of the chromatogram peak and assigning the candidate chemical structure to the chromatogram peak within the second group having the score indicating the greatest correlation, wherein the score is calculated based on identifying one or more m / z peak(s) in the associated mass spectral data corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure.
19. The method of any one of claims 16 to 18, wherein the score is calculated based on identifying, in the associated mass spectral data, a number and / or intensity of the one or more m / z peak(s) corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure.
20. The method of any one of claims 16 to 19, wherein the score is calculated based on a ratio of the sum of intensities of the one or more m / z peak(s) corresponding to the m / z ratio(s) of the unique fragment(s) of the candidate chemical structure to the sum of intensities of all m / z peaks in the associated mass spectral data.
21. The method of any one of claims 16 to 20, wherein the score is calculated based on a ratio of a number of the one or more m / z peak(s) corresponding to the m / z ratio(s) of unique fragment(s) of the candidate chemical structure to a number of all m / z peaks in the associated mass spectral data.
22. A computer program comprising instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 21.
23. An apparatus, comprising: a chromatography instrument configured to obtain the chromatogram, a mass spectrometer configured to obtain the associated mass spectral data for each chromatogram peak; and a data analysis system configured to carry out the method of any one of claims 1 to 21 using the chromatogram obtained by the chromatography instrument and associated mass spectral data obtained by the mass spectrometer.
Citation Information
Patent Citations
A precise identification method for unknown volatile components in scented stationery
CN114088857B
Mass spectrometer and method of analyzing isomers
US20050258355A1
High Mass Accuracy Filtering for Improved Spectral Matching of High-Resolution Gas Chromatography-Mass Spectrometry Data Against Unit-Resolution Reference Databases
US20150340216A1