Method for analysis of a chromatogram

A machine learning model predicts retention parameters from structural formulas to order and assign chromatogram peaks, addressing the challenge of isomer differentiation in derivative mixtures, enhancing peak identification precision.

WO2025261606A1PCT designated stage Publication Date: 2025-12-26HIGHCHEM SRO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/067459
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing chromatography techniques struggle to unambiguously assign peaks in a chromatogram of a sample containing a mixture of derivatives, especially isomers, due to their similar masses and varying degrees of derivatization, making mass spectroscopy unreliable.

Method used

A method using a machine learning model to predict retention parameters based on structural formulas of candidate compounds, ordering them by retention time, and correlating peak order with compound order to assign peaks accurately.

Benefits of technology

Enables peak assignment in chromatograms without mass spectroscopy, effectively distinguishing isomers and derivatives, improving peak identification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024067459_26122025_PF_FP_ABST
    Figure EP2024067459_26122025_PF_FP_ABST
Patent Text Reader

Abstract

In accordance with the claimed invention, there is provided a method for analysis of a chromatogram. The method comprising obtaining a chromatogram of a sample comprising a mixture of compounds, providing one or more parameter(s) obtained from a structural formula of each of a plurality of candidate compounds to a machine learning model and generating using the machine learning model a retention parameter indicative of retention time for each of the candidate compounds as output, ordering the candidate compounds according to their retention parameter and assigning peaks of the chromatogram to the candidate compounds based on correlating an order of the peaks according to their retention time to the order of the candidate compounds according to their retention parameter such that the candidate compound having the retention parameter indicative of the shortest retention time is assigned to the peak with the shortest retention time. The mixture of compounds comprises a plurality of derivatives of a known precursor and a plurality of the candidate compounds are predicted derivatives of the known precursor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for Analysis of a Chromatogram

[0002] Field

[0003] The claimed invention relates to a method for analysis of a chromatogram. Particularly, a method that enables assigning of peaks of a chromatogram of a sample containing a mixture of derivatives of a precursor.

[0004] Background

[0005] Derivatization is commonly employed in gas chromatography-mass spectrometry (GC-MS) and liquid chromatography-mass spectrometry (LC-MS). As is well known in the art, Derivatization is the process of chemically altering a compound by reacting the compound with a derivatization agent. Derivatization of a compound is often performed before input to the chromatography column to improve, for example, the chromatographic separation and / or thermal stability of the compound. Derivatization can involve, for example, silylation, alkylation or acylation reactions. During derivatization, the derivatization agent reacts with a functional group in the compound, for example active hydrogens in the case of silylation. Such functional groups may be referred to as active functional groups. If the compound contains multiple functional groups, then derivatization of the compound can result in a mixture of derivatives where the derivatized functional group may be in a different position in each derivative. Furthermore, the different derivatives may exhibit varying degrees of derivatisation. Typically, an excess of derivatization agent is used in the derivatization reaction so that the predominant derivative generated has all active functional groups derivatised. However, there may still be other partially derivatised compounds present. In a chromatogram of a sample produced from derivatization of a compound, in addition to the dominant peak corresponding to the fully derivatized compound, there may be less intense peaks corresponding to partially derivatized compounds. There may also be a peak corresponding to the underivatized compound.

[0006] As a number of different derivatizations are possible for a compound with multiple functional groups, it may be difficult to determine which derivative should be assigned to which peak of a chromatogram of a sample of a derivatised compound. This is especially the case for a sample comprising derivatives that are also isomers, since such isomers would have the same mass but differ in their structural formula. Therefore, known techniques of using mass spectroscopy to assign peaks of a chromatogram cannot readily be relied on for unambiguously assigning peaks of a chromatogram of a sample where the sample has a mixture of isomers.

[0007] Summary

[0008] Against this background, there is provided a method for analysis of a chromatogram, comprising: obtaining a chromatogram of a sample comprising a mixture of compounds, wherein the mixture of compounds comprises a plurality of derivatives of a known precursor; providing one or more parameters obtained from a structural formula of each of a plurality of candidate compounds to a machine learning model and generating using the machine learning model a retention parameter indicative of retention time for each of the candidate compounds as output, wherein a plurality of the candidate compounds are predicted derivatives of the known precursor; ordering the candidate compounds according to their retention parameter; and assigning peaks of the chromatogram to the candidate compounds based on correlating an order of the peaks according to their retention time to the order of the candidate compounds according to their retention parameter such that the candidate compound having the retention parameter indicative of the shortest (lowest) retention time is assigned to the peak with the shortest (lowest) retention time.

[0009] A computer program comprising instructions, that when executed by a processor, cause the processor to perform any method herein disclosed is also provided. An apparatus comprising: a chromatography instrument configured to obtain the chromatogram; and a data analysis system configured to carry out any method herein disclosed using the chromatogram obtained by the chromatography instrument is also provided.

[0010] The method of the claimed invention enables peaks of a chromatogram of a sample comprising a mixture of derivatives of a known precursor to be identified without performing mass spectroscopy. This method may be applicable to chromatograms obtained using gas chromatography or using liquid chromatography.

[0011] The method is particularly advantageous for enabling assignment of peaks of a chromatogram of a sample of a derivatised compound where some of the derivatives are isomers. In known methods of assignment of peaks of a chromatogram, a mass spectra associated with each peak would be used to ascertain which peak to assign to which compound. However, such known techniques do not readily facilitate the identification of isomers that have the same mass but different structures. The method of the claimed invention does enable assignment of peaks to their corresponding isomers without needing to fragment the compounds within the sample or perform mass spectroscopy.

[0012] Derivatives of the known precursor are compounds resulting from derivatisation of the known precursor. Derivatization is the chemical reaction between the known precursor and a derivatisation agent to form the derivatives. Functional groups of the known precursor react with the derivatization agent to form derivatized groups. The known precursor may comprises a plurality of functional groups. The functional groups may be referred to as active functional groups. Some examples of functional groups include: alcohols, alkenes, alkynes, amines, carboxylic acids, aldehydes, ketones, esters, active hydrogens and ethers.

[0013] If the known precursor has a plurality of functional groups, then the derivatization reaction can result in a mixture of derivatives with different degrees of derivatization and various positions of the derivatized group within the structure. Some of those derivatives may have the same degree of derivatization but different positions of the derivatized group within the structure thereby resulting in isomers (compounds having the same molecular formula but different structural formula).

[0014] Derivatisation and the use of derivatisation agents is well known in the art. In known chromatography techniques, derivatization is often performed to input to the chromatography column improve, for example, the chromatographic separation and / or thermal stability of the compound.

[0015] Retention parameter used herein refers to a parameter indicative of retention time. The retention parameter may be, for example, an estimated retention index. Alternatively, the retention parameter may be, for example, an estimated retention time. As is known in the art, the retention time of a sample is a measure of the time between injection of the sample and elution of the sample from the chromatography column (from injection to detection).

[0016] Retention times vary with the individual chromatographic system, with regard to column length, diameter, temperature, flow rates etc. Retention indices (also referred to as retention indexes) are well known in the art. Retention indices are indicative of the retention time but independent of measurement condition. Retention indices normalise the retention time of a compound by the retention time of a predetermined standard substance. For example, the Kovats retention indices normalise the retention time of the sample to the retention time of predefined standard (usually n-alkanes) and column polarity. The method may also involve obtaining the sample by performing derivatization of the known precursor. The derivatization may be performed using a derivatization agent, as would be known in the art.

[0017] The known precursor may be any compound where the structural formula of the precursor is known. It is not necessary for the molecular formula of the precursor to be known to generate the derivatives in silico. However, optionally, both the structural and molecular formula of the precursor may be known.

[0018] The known precursor may be referred to as a known compound. The known precursor may comprises a plurality of functional groups. The functional groups may react with the derivatization agent. The functional groups may be referred to as active functional groups. Some examples of functional groups include: alcohols, alkenes, alkynes, amines, carboxylic acids, aldehydes, ketones, esters, active hydrogens and ethers.

[0019] A plurality of the candidate compounds are predicted derivatives of the known precursor. In one example, all of the candidate compounds may be predicted derivatives of the known precursor. In an alternative example, one of the candidate compounds may be the known precursor (i.e. the underivatized form of the precursor) and the rest of the candidate compounds may be predicted derivatives of the known precursor. The structural formula of each of the predicted derivatives may be determined by generating derivatives of the known compound by simulating the chemical reaction of the derivatization in-silico (i.e. by computer modelling or simulation).

[0020] For each candidate compound, one or more parameters obtained from the respective structural formula are input to the machine learning model to generate the retention parameter for that candidate compound. In other words, for each candidate compound, the retention parameter is generated based on the one or more parameters obtained from the structural formula.

[0021] The one or more parameters may comprise the structural formula of the candidate compound. The one or more parameters may be the structural formula of the candidate compound.

[0022] The one or more parameters may comprise atomic attribute(s) of atom(s) of the structural formula and / or bond attributes of pair(s) of atoms, which may be bonded together or unbonded, of the structural formula. The one or more parameters may comprise atomic attribute(s) of one or more atom(s) of the structural formula and / or bond attributes of one or more pair(s) of atoms of the structural formula. The one or more parameters may comprise atomic attribute(s) of each atom of the structural formula and / or bond attributes of each bond of the structural formula. “Atomic attributes” refers to properties of the atoms within the structural formula. “Bond attributes” refers to properties of pair of atoms, which may be bonded together, within the structural formula.

[0023] The atomic attributes may comprise one or more of hybridisation of the atom, atom type, formal charge of the atom, hydrogen bonding type of the atom, whether the atom is part of an aromatic ring, degree of the atom, number of hydrogens the atom is connected to, chirality of the atom and / or partial charge of the atom. Hydrogen bonding of the atom being whether the atom is a hydrogen bond donor or acceptor. These atomic attributes may be derived from the structural formula. The atom type, hybridisation, whether the atom belongs to an aromatic ring, the degree of the atom, the number of hydrogens the atom is connected to and the chirality may each be individually represented by a one-hot vector when input to the model.

[0024] The bond attributes may comprise one or more of the bond type of the pair of atoms (e.g. no bond, single bond, double bond, triple bond, aromatic bond etc.), whether the bond is between atoms in the same ring, whether the bond is conjugated or not, the stereoconfiguration of the bond. These bond attributes may be derived from the structural formula. These parameters may be represented by a one-hot vector when input to the model.

[0025] The one or more parameter(s) obtained from the structural formula and input to the model may also comprise one or more global attribute(s) of the structural formula. Global attributes being properties of the structural formula as a whole. Global attributes may comprise number of atoms in the structural formula, molecular mass of the structural formula.

[0026] Preferably, the number of candidate compounds is the same as the number of peaks in the chromatogram. This enables a direct correlation between the order of the predicted derivatives according to their retention parameter and the order of the peaks according to their retention time.

[0027] The method may comprise comprising calculating a regression between the retention time of the peaks and the retention parameter of the corresponding assigned candidate compounds. In other words, for each peak, calculating the regression between the retention time of the peak and the retention parameter of the candidate compound assigned to that peak.

[0028] The method may further comprise obtaining the chromatogram by performing gas or liquid chromatography of the sample. Gas and liquid chromatography techniques are well known in the art.

[0029] Brief Description of the Drawings

[0030] Figure 1 shows a flow diagram indicating steps of a method in accordance with an embodiment of the disclosure.

[0031] Figure 2(a) is an exemplary chromatogram.

[0032] Figure 2(b) is a plot indicating the correlation between the retention time (t«) of peaks of the exemplary chromatogram of Figure 2(a) and the retention parameter of the corresponding candidate compound as determined by the machine learning model in accordance with the method. The retention parameter is an estimated retention index (Rl) in this exemplary embodiment.

[0033] Figure 3(a) is a plot of retention time (tps) against retention parameter, the retention parameter in this case being a retention index (Rl). The plot indicating a theoretical linear correlation between retention time (t«) of peaks of a chromatogram and a theoretical retention index (Rl). The theoretical linear relationship being represented as a straight line on the plot, referred to herein as a theoretical best fit line.

[0034] Figures 3(b) and 3(c) are exemplary plots of retention time (tps) against retention parameter used to explain how the theoretical best fit line of Figure 3(a) can be used to identify inaccurate retention indices generated using the machine learning model for quality control of the machine learning model.

[0035] Figure 4 is a first exemplary implementation of the method of Figure 1 where the sample used to obtain the chromatogram comprises derivatives of tryptophan.

[0036] Figure 5 is a second exemplary implementation of the method of Figure 1 where the sample used to obtain the chromatogram comprises cimifugin. Detailed Description

[0037] The method of the claimed invention is now described in accordance with the flow diagram of Figure 1.

[0038] Step 101

[0039] As set out in step 101 , first a chromatogram of the sample is obtained. The chromatogram may be pre-stored in and obtained from a database or memory. The chromatogram may have been produced by gas chromatography or liquid chromatography of the sample. In one embodiment, the method may comprise the step of performing gas chromatography or liquid chromatography to obtain the chromatogram.

[0040] In chromatography, the sample to be analysed is supplied to a chromatography instrument that has a column and a detector. The sample is carried by a mobile phase through the column. The column has a stationary phase that may contain a plurality of particles.

[0041] Compounds of the sample elute at different rates from the column according to their degree of interaction with the stationary phase. The compounds of the sample eluting from the column are detected by a detector.

[0042] A gas chromatography instrument employs a gas chromatography column and, in use, a sample is vapourised and injected into the gas chromatography column together with a carrier gas as the mobile phase. The carrier gas is typically an inert gas.

[0043] A liquid chromatography instrument employs a liquid chromatography column and, in use, a sample is dissolved in a liquid mobile phase and injected into the liquid chromatography column together with the liquid mobile phase. The mobile phase is pressurised in high performance liquid chromatography.

[0044] Detectors, in particular mass spectrometry detectors, used in chromatography instruments are well known. Examples of liquid chromatography detectors include absorbance detectors, fluorescence detectors, evaporative light scattering detectors, fluorescence / chemiluminescence detectors, electrochemical and refractive index detectors. Examples of gas chromatography detectors include flame ionization detectors, thermal conductivity detectors, barrier discharge ionization detectors, electron capture detectors, sulphur chemiluminescence detectors, nitrogen phosphorus detectors, thermal conductivity detectors and pulsed discharge detectors. A chromatogram may be produced by measuring the quantity of sample molecules which elute from the column over time using the detector. Sample compounds (compounds within the sample) which elute from the column will be detected as a peak above a baseline measurement on the chromatogram. Where different compounds within a sample have different elution rates, a plurality of peaks on the chromatograph may be detected. Ideally, individual sample peaks are separated in time from other peaks in the chromatogram such that different compounds do not interfere with each other.

[0045] Typically a chromatographic peak has a Gaussian shaped profile, or can be assumed to have a Gaussian shaped profile.

[0046] A chromatogram is a plot of detector response (detector signal intensity) against retention time. The retention time of a compound is generally measured as the period of time between injection of sample into the column and the relative intensity peak maximum after chromatographic separation by the column. The retention time depends on the compound and also on the measurement conditions, such as the composition of the stationary phase, the pressure of the column, the temperature of the column and the dimensions of the column, as is known in the art.

[0047] Chromatography instruments and the resulting chromatograms produced are well known in the art and so will not be described in any further detail. This outline is simply to provide context for the method of the claimed invention that involves analysis of a chromatogram.

[0048] The chromatogram generated by performing chromatography may be output to a data analysis system that performs the method steps set out in Figure 1 .

[0049] As set out in step 101 , the chromatogram is of a sample comprising a mixture of compounds. A plurality of those compounds are derivatives of a known precursor. The known precursor may be any compound where the structural formula of the precursor is known.

[0050] It is known that the sample contains a plurality of derivatives of a known precursor. However, it is not known which derivative should be assigned to which peak of the chromatogram. In some embodiments, the sample may contain further compounds , such as the known precursor in its underivatized form. In such embodiments, it is known that the sample contains both the known precursor in its underivatized form and derivatives of the known precursor but it is not known which peak of the chromatogram should be assigned to each of the derivatives or to the underivatized precursor. The sample may consist of (that is, comprise only) derivatives of the known precursor. Alternatively, the sample may consist of (that is, comprise only) derivatives of the known precursor and the known precursor in its underivatized form.

[0051] As discussed above, the chromatogram of the sample may be pre-stored in a memory or database. The chromatogram may be labelled so that the skilled person knows the structural formula of the precursor used to produce the sample.

[0052] Alternatively, the method may further comprise the step of obtaining the sample by performing derivatisation of the known precursor by reacting the known precursor with a derivatisation agent and may further comprise the step of performing chromatography of the sample to obtain the chromatogram. It may be known that the sample contains both the underivatized precursor and the derivatives if, for example, it is known that derivatisation agent was not used in excess to produce the sample. Similarly, it may be known that the sample only contains derivatives of the precursor based on knowledge that the derivatisation agent was used in excess.

[0053] The derivatives within the sample are compounds resulting from derivatization of the known precursor. Derivatization is the chemical reaction between the known precursor and a derivatisation agent to form the derivatives. The functional groups of the known precursor react with the derivatization agent to form derivatized groups.

[0054] The known precursor may comprise a plurality of functional groups. The functional groups may be referred to as active functional groups. Some examples of functional groups include: alcohols, alkenes, alkynes, amines, carboxylic acids, aldehydes, ketones, esters, active hydrogens and ethers.

[0055] If the known compound has a plurality of functional groups, then the derivatization reaction can result in a mixture of derivatives with different degrees of derivatization and various positions of the derivatized group within the structure. Some of those derivatives may have the same degree of derivatization but different positions of the derivatized group within the structure thereby resulting in isomers (compounds having the same molecular formula but different structural formula). Derivatisation and the use of derivatisation agents is well known in the art. In known chromatography techniques, derivatization is often performed to input to the chromatography column improve, for example, the chromatographic separation and / or thermal stability of the compound. Derivatization can involve, for example, silylation, aklylation or acylation reactions. During derivitization, the derivitization agent reacts with an active functional group in the compound, for example active hydrogens in the case of silylation. Derivatisation agents include, for example, N,0-bis(trimethylsilyl)trifluoroacetamide, 1 ,3-bis(chloromethyl)-1 ,1 ,3,3-tetramethyldisilazane, dimethylchlorosilane, N- heptafluorobutyrylimidazole.

[0056] Step 102

[0057] In Step 102, for each candidate compound, one or more parameters obtained from the structural formula of the candidate compound are input to a machine learning model to generate a retention parameter for the candidate compound.

[0058] The candidate compounds are compounds that are predicted to be those compounds in the sample and so correspond to the peaks of the chromatogram. As discussed above, the sample comprises derivatives of the known precursor. Consequently, a plurality of the candidate compounds are predicted derivatives of the known precursor.

[0059] In one embodiment, all of the candidate compounds may be predicted derivatives of the known precursor. Alternatively, one of the candidate compounds may be the known precursor (i.e. the known precursor in its underivatized form) and the rest of the candidate compounds may be predicted derivatives of the known precursor. For example, in embodiments where the sample has been produced using an excess of derivatisation agent, it may be expected that the precursor is fully derivatized and so the sample and the candidate compounds may not include the known precursor. However, in embodiments, where an excess of derivatisation agent has not been used, then it may be expected that the precursor is not fully derivatised and so the sample and the candidate compounds may include the known precursor.

[0060] The structural formula of each of the predicted derivatives may be determined by simulating the derivatization reaction between the known precursor and derivatization agent using computer modelling (i.e. in-silico). It is possible to use a software package to simulate chemical reactions in-silico based on input of the 2D structural formula of the reagents to be reacted (in this case the known precursor and the derivatisation agent). An exemplary software package is RDKit that can be used to simulate the derivatization reaction between a known precursor and derivatisation agent using computer modelling. The chemical reactions may be constructed using, for example, a SMARTS-based naming format, similar to Daylight’s Reaction SMILES or SMIRKS as is known in the art. The predicted derivatives may be stored in a database in association with the known precursor and derivatization agent.

[0061] Alternatively, the structural formula of each of the predicted derivatives may have been previously determined by generating derivatives of the known compound by experimentation. The structural formula of each of the predicted derivatives of the known compound may be pre-stored in a database in association with the known precursor and derivatization agent and obtained from that database.

[0062] In step 102 of the method, one or more parameters obtained from the structural formula of each candidate compound is input to a machine learning model to generate a retention parameter for each candidate compound.

[0063] The structural formula may be provided using any known chemical naming formats, for example, SMARTS, SMILES (simplified molecular-input line-entry system), Isomeric SMILES, SMIRKS, InChi (International Union of Pure and Applied Chemistry), CML, SDF- files or MOL-file. These formats are all known in the art.

[0064] The retention parameter is typically an estimated retention index. The retention index may be a Kovats Retention Index. Kovats Retention Indices are well known in the art. However, the invention may be implemented with other types of retention parameters, such as an estimated retention time.

[0065] Machine learning models that predict retention parameters based on one or more parameters obtained from a structural formula are well known. The machine learning model may be the MEGNet model (MatErial Graph Network), where one or more parameter(s) obtained from the structural formula of the compound is input and a retention parameter is output from the model. The structural formula may be provided as a MOL-file. The one or more parameters input from the model may be derived from or obtained from the MOL-file may be input to and used by the model to generate the retention parameters.

[0066] Parameter(s) obtained from the structural formula and input to the machine learning model may be the structural formula. Indeed, the structural formula may be mapped directly to the machine learning model. The structural formula may be provided in one of the known chemical formats set out above.

[0067] Parameter(s) obtained from the structural formula and input to the machine learning model may be atomic attributes of one or more atom(s) within the structural formula and / or bond attribute(s) of one or more pair(s) of atoms, which may or may not be bonded together,) within the structural formula. Parameter(s) obtained from the structural formula and input to the machine learning model may be atomic attributes of each atom within the structural formula and / or bond attribute(s) of each pair of atom(s) within the structural formula.

[0068] Atomic attributes may be referred to herein as atomic properties i.e. properties of an atom within the structural formula. Bond attributes may be referred to herein as bond properties i.e. properties of a bond within the structural formula.

[0069] The atomic attributes may comprise one or more of hybridisation of the atom, atom type, formal charge of the atom, hydrogen bonding type of the atom, whether the atom is part of an aromatic ring, degree of the atom, number of hydrogens the atom is connected to, chirality of the atom and / or partial charge of the atom. Hydrogen bonding of the atom being whether the atom is a hydrogen bond donor or acceptor. These atomic attributes may be derived from the structural formula. For example, where the structural formula is provided as a MOL-file, these atomic attributes may be derived from the MOL-file. The atom type (e.g. C, N,O,F etc.), hybridisation (e.g. sp, sp2, sp3), whether the atom belongs to an aromatic ring, the degree of the atom (number of directly bonded neighbours), the number of hydrogens the atom is connected to and the chirality (e.g. R or S) may each be individually represented by a one-hot vector when input to the model. In one embodiment, the atomic attributes comprise, for each atom, hybridisation of the atom, atom type, formal charge of the atom, hydrogen bonding type of the atom, whether the atom is part of an aromatic ring, degree of the atom, number of hydrogens the atom is connected to.

[0070] The bond attributes may comprise one or more of the bond type (e.g. no bond, single, double, triple, aromatic), whether the bond is between atoms in the same ring, whether the bond is conjugated or not and / or the stereo-configuration of the bond. These bond attributes may be derived from the structural formula. For example, where the structural formula is provided as a MOL-file, these bond attributes may be derived from the MOL-file. These parameters may be represented by a one-hot vector when input to the model.

[0071] Such parameters are described here: https: / / deepchem.readthedocs.io / en / latest / api reference / featurizers.html#molqraphconvfeat urizer, which is incorporated herein by reference, and in “Predicting Kovats Retention Indices

[0072] Using Graph Neural Networks” Chen Qu et. Al., Journal of Chromatography A, 1646 (2021 ), which is incorporated herein by reference.

[0073] The one or more parameter(s) obtained from the structural formula and input to the model may also comprise global attributes of the structural formula. Global attributes being properties of the structural formula as a whole. Global attributes may comprise number of atoms in the structural formula, molecular mass of the structural formula.

[0074] The model may use graph neural networks. The model may be trained using parameter(s) obtained from structural formulas and corresponding retention parameters, such as retention indexes or retention times as the training data. The parameter(s) obtained from structural formulas used to train the model may be those discussed above. The training data may comprise Kovats retention indices from the NIST database, or experimental retention times, or retention times from the METLINE small molecule retention time (SMRT) database. For an embodiment concerning predicting retention time for a gas chromatogram, the training data may comprise one or more of: column type, column length, column diameter, column temperature, temperature gradient within the column, gas pressure within the column and flow rate within the column. For an embodiment concerning predicting retention time for a liquid chromatogram, the training data may comprise one or more of column type, column length, column diameter, column temperature, flow rate within the column, mobile phase composition and gradient of the mobile phase composition.

[0075] Once the machine learning model has been trained, one or more parameter(s) obtained from the 2D structural formula of each of the candidate compounds can be input to the model described above and a corresponding retention parameter can be generated. The retention parameter generated using the model may be a retention index or retention time. If the data used to train the model comprises retention indices, then the retention parameter generated using the model may be a retention index. If the data used to train the model comprises retention times, then the retention parameter generated using the model may be a retention time. In one particular example, if the data used to train the model employs the Kovats retention index, then the retention index generated by the model would be a Kovats retention index.

[0076] The one or more parameter(s) obtained from the 2D structural formula model may be the 2D structural formula of the candidate compound. In other words, the 2D structural formula of the candidate compound may be input to the model using any known chemical naming formats, for example, SMARTS, SMILES (simplified molecular-input line-entry system), Isomeric SMILES, SMIRKS, InChi (International Union of Pure and Applied Chemistry), CML, SDF-files or MOL-file. These naming systems are all known in the art. In particular, the MOL-format is typically preferred.

[0077] Exemplary Machine Learning Model

[0078] An exemplary machine learning model that can be employed in the method of the invention to calculate the retention parameter based on parameters obtained from a structural formula of a compound is described in “C. Chen, W. Ye, Y. Zuo, C. Zheng, S.P. Ong, Graph networks as a universal machine learning framework for molecules and crystals, Chem. Mater. 31 (9) (2019) 3564 - 3572”, which is incorporated herein by reference. This model is provided in deepchem, as set out here https: / / deepchem.readthedocs.io / en / latest / api reference / models.html#megnetmodel.

[0079] The use of this model to predict retention indices is described in “Predicting Kovats Retention Indices Using Graph Neural Networks” Chen Qu et. Al., Journal of Chromatography A, 1646 (2021), which is incorporated herein by reference. This paper therefore exemplifies how this model can be used in the claimed invention determine a retention parameter (in this case retention indexes) based on parameters obtained from the structural formula. As described in “Predicting Kovats Retention Indices Using Graph Neural Networks”, the model predicts the retention index of a compound based on parameters obtained from the 2D structural formula of the compound input as a MOL-file. The model employed is the MEGNet model. The model is trained using parameters obtained from structural formulas and Kovats retention indices from the NIST database.

[0080] Training the Exemplary Machine Learning Model

[0081] In this exemplary machine learning model described in “Predicting Kovats Retention Indices Using Graph Neural Networks” that can be employed in the claimed invention, the training data is obtained from the NIST Standard Reference Database. The training data input to the model are an experimental value for the Kovats retention index and parameters (atomic (node) attributes, bond (edge) attributes and global (state) attributes) obtained from a 2D representation of the structural formula . The 2D representation of the structural formula can be in a number of formats, such as Isomeric SMILES, InChi (International Union of Pure and Applied Chemistry), CML, SDF-files or MOL-files. These formats are all known in the art. These formats all capture the connectivity within the compound. In this exemplary implementation, the 2D representation of the structural formula is provided as a MOL-file.

[0082] The MEGNet model employs a graph neural network. The structural formula of the compound may be mapped directly onto the graph due to the direct correspondence between them. The nodes in the graph correspond to atom locations and edges in the graph correspond to pairs of atoms (not necessarily chemically bonded together).

[0083] The nodes have atomic attributes / properties assigned thereto and the edges have bond attributes / properties assigned thereto. Global attributes / properties of the compound not specific to certain atoms or bonds e.g. the molecular mass are also incorporated into the graph.

[0084] The atomic attributes include atomic number, hybridisation and formal charge on the atom. The hybridization may be calculated from information in the MOL-file with use of RDKit, as is known in the art. The edge attributes include bond type, graph distance and whether atoms are in the same ring. The global attributes include the number of non-hydrogen atoms in the structural formula, the molecular mass of the structural formula divided by the number of non-hydrogen atoms, and the number of chemical bonds in the structural formula divided by the number of non-hydrogen atoms.

[0085] The initial graph of the model is represented by the set of atomic attributes, bond attributes and global attributes. The attributes, and so the graph, is then successively updated to form a graph neural network layer. Firstly, the bond attributes are updated using attributes from the bond, from neighbouring atoms and from the global attributes using a bond update function and a concatenation operator. Secondly, the atomic attributes are updated using attributes from itself, the bonds connected to the atom and the global attributes using an atom update function and a concatenation operator. Thirdly, the global attributes are updated using information from itself and all atoms and bonds based on a global state update function and a concatenation operator. The update functions may be multilayer perceptrons with a rectified linear unit activation function. Each MEGNet block may contain two dense layers of multi-layer perceptrons and the graph neural network layer. The two dense layers may be before the graph neural network to pre- process the input. The MEGNet model may be used with multiple MEGNet blocks. In particular, three MEGNet blocks may be employed.

[0086] After the MEGNet blocks, there is a readout operator that reduces the output graph to a scalar or a vector. In this example, the “Set2set” model may be applied on the atomic and bond attributes to reduce these vectors into one vector. Subsequently there is a concatenation step and two further densely-connected multi-layer perceptrons.

[0087] To fit the model, an Adam optimizer may be used and the mean absolute error may be employed as the loss function. The Kovats retention indices taken from the NIST library may be used as target values. These retention indices may be normalised where the normalised value is the experimental from the database minus the mean of the experimental Kovats retention indices used in the model divided by the standard deviation of the experimental Kovats retention indices used in the model.

[0088] Alternative models

[0089] The exemplary machine learning model described above uses Kovats retention indices from the NIST database in the training data. When such a model is used by inputting a one or more parameter(s) obtained from a structural formula of a compound, the resulting retention parameter that is generated is an estimated Kovats retention index. However, other types of retention parameters can be used in training the model and consequently other types of retention parameters may be output from the machine learning model.

[0090] For example, the model may be trained based on the one or more parameter(s) obtained from structural formulas of compounds and their experimentally measured retention times. By way of further example, the model may be trained based on one or more parameter(s) obtained from structural formulas of compounds and the retention times obtained from the METLIN small molecule retention time (SMRT) dataset. This may be particularly applicable for methods involving chromatograms obtained by liquid chromatography. The retention parameter that would be output from such a model would be an estimated retention time.

[0091] An example of a model that employs graph neural networks to generate retention times for liquid chromatography based on molecular structure is described in “Prediction of Liquid Chromatographic Retention Time with Graph Neural Networks to Assist in Small Molecule Identification”, Qion Yang et al., Analytical Chemistry, 2021 , 93, 2200-2206, which is hereby incorporated by reference. The training data input to such a model may be one or more parameter(s) obtained from the structural formula and the corresponding retention time for liquid chromatography taken from the METLIN small molecule retention time (SMRT) dataset.

[0092] Step 103

[0093] Once a retention index for all of the candidate compounds has been generated, the candidate compounds are ordered according to their retention parameter.

[0094] As discussed above, the retention parameter is indicative of retention time. The retention parameter may be an estimated retention time or an estimated retention index.

[0095] As is known in the art, the retention time of a sample is a measure of the time between injection of the sample and elution of the sample from the chromatography column (from injection to detection). Retention times vary with the individual chromatographic system, with regard to column length, diameter, temperature, flow rates etc. Retention index is a dimensionless parameter indicative of retention time. Retention indices are indicative of the retention time but independent of measurement condition. Retention indices normalise the retention time of a compound by the retention time of a predetermined standard set, usually n-alkanes. For example, the Kovats retention indices normalise the retention time of the sample to the retention time of standard compounds, usually n-alkanes, at a given column polarity. Retention indices are typically employed in gas chromatography.

[0096] In step 103, the candidate compounds are ordered according to their retention parameter ideally from the retention parameter indicative of the shortest retention time to the retention parameter indicative of the longest retention time.

[0097] Peaks appear on a chromatogram in order of their retention time, which is on the x-axis of the chromatogram. The order of the candidate compounds according to their retention parameters directly corresponds with the order of the corresponding peaks so that the peak having the shortest (lowest) retention time corresponds to the candidate compound having the retention parameter indicative of the shortest (lowest) retention time.

[0098] Step 104 Step 104 of the method involves assigning the peaks of the chromatogram to the candidate compounds based on correlating an order of the peaks according to their retention time to the order of the candidate compounds according to their retention parameter such that the candidate compound having the retention parameter indicative of the shortest retention time is assigned to the peak with the shortest retention time. The candidate compounds are assigned to the peaks of the chromatogram so that the order of the candidate compounds according to their retention parameter is the same as the order of the corresponding peaks ordered according to their retention time.

[0099] Optionally, a smaller retention parameter may be indicative of a shorter retention time. In such an embodiment, the candidate compounds may be ordered from smallest to greatest retention parameters. The peaks from left to right of the chromatogram (i.e. from shortest to longest retention times) may be assigned to the candidate compounds in order from smallest to greatest retention parameters.

[0100] Preferably, the number of candidate compounds is the same as the number of peaks in the chromatogram. This enables a direct correlation between the order of the candidate compounds according to their retention parameter and the order of the peaks according to their retention time. In one embodiment, the candidate compounds may consist of (that is, comprise only) the known precursor and predicted derivatives of the known precursor. In such an embodiment, the number of peaks may be the same as the total number of predicted derivatives plus one for the known precursor. In another embodiment, the candidate compounds may consist of (that is, comprise only) the predicted derivatives of the known precursor. In such an embodiment, the number of peaks may be the same as the total number of predicted derivatives.

[0101] By way of a further method step, not shown in Figure 1 , once the peaks have been assigned to the corresponding candidate compound, the retention time of the peak and the retention parameter of the corresponding candidate compound calculated by the machine learning model can be correlated.

[0102] The regression between the retention times of the peaks of the chromatogram and the retention parameters of the corresponding candidate compound may be calculated. For example, the retention time may be plotted against the retention parameter of the corresponding candidate compound and a line of best fit identified. The statistics of the regression can be calculated and used as a quality measurement of the model. The inventors have found that the regression may be logarithmic for isothermal chromatography (chromatography performed with the column at a constant temperature) and linear for thermal gradient chromatography (chromatography performed with the column having a thermal gradient along it).

[0103] An exemplary plot is shown in Figure 2(b) that corresponds to the chromatogram of Figure 2(a). In the plot of Figure 2(b), the x-axis is the retention time (t«) of the peaks of the chromatogram and the y-axis is the retention parameter, which in this case is retention index (Rl), as calculated by the machine learning model. The four points plotted correspond to the four peaks of the chromatogram. Each point has an x-co-ordinates that is the retention time of the peak and a y-coordinate that is the retention index calculated for the corresponding candidate compound. In Figure 2(b), the plot indicates a linear correlation i.e. linear regression.

[0104] Figure 3 exemplifies how identification of a theoretical linear relationship between retention index and retention time can be used to indicate a quality measurement of the model. The plot of Figure 3(a) depicts a theoretical (perfect, i.e. ideal) linear correlation between the retention time on the x-axis and the retention parameter (which in this case is retention index). The theoretical linear relationship is represented as a straight line on the plot (referred to herein as the theoretical best fit line). The theoretical best fit line may be calculated by linear regression.

[0105] Data points of the retention time (tps) of peaks of a chromatogram and the retention parameter, in this case retention index, of the corresponding candidate compound as determined by the machine learning model in accordance with the method can then be plotted on a plot including this theoretical best fit line, as shown in Figures 3(b) and 3(c).

[0106] If a data point of the retention time of a chromatogram peak and the retention parameter, in this case retention index, of the corresponding candidate compound as determined by the machine learning model does not significantly deviate from this theoretical best fit line, then it indicates that the retention index is likely to be accurate for that chromatogram peak and the machine learning model is likely to be of good quality, as shown in Figure 3(b).

[0107] If a data point of the retention time of a chromatogram peak and the retention parameter, in this case retention index, of the corresponding candidate compound as determined by the machine learning model does significantly deviate from this theoretical best fit line, then it indicates that the retention index is likely to be inaccurate for that chromatogram peak and the machine learning model is not likely to be of good quality. For example, the data point having a retention index of 2400 and a retention time of approximately 8.58 minutes significantly deviates from the theoretical best fit line thereby indicating that the retention index is unlikely to be accurate for the chromatogram peak corresponding to the retention time of 8.58 minutes.

[0108] The data points in Figure 3 are all exemplary values rather than data points actually calculated using the machine learning model.

[0109] Exemplary implementations

[0110] Figure 4 sets out a first exemplary implementation of the method of the claimed invention where each of steps 101 , 102, 103 and 104 are shown as separate sections of the figure divided by dashed lines.

[0111] In Figure 4, the known precursor is tryptophan having the structural formula shown in the step 101 section of Figure 4. As can be seen, tryptophan has multiple functional groups. For tryptophan, those functional groups are active hydrogens that are reactive with a derivatisation agent to form the derivatives. In this case, it is known that the derivatization reaction is silylation. It may be known that the derivatization agent has been reacted with tryptophan in excess to obtain the sample so that the sample does not contain any underivatized forms of tryptophan.

[0112] Performing chromatography to obtain a chromatogram of the sample would be well known in the art. As shown in the step 101 section of Figure 4, the chromatogram has two peaks, each peak corresponding to one of the derivatives of tryptophan.

[0113] The skilled person with knowledge of the known precursor (tryptophan) and knowledge of the derivatization agent (BSTFA (N,Obis( trimethylsilyl)trifluoroacetamide)) would be able to obtain the candidate compounds, which in this case are predicted derivatives of tryptophan where the derivatization is silylation.

[0114] In this example, the structural formula of each of the predicted derivatives of tryptophan is determined by simulating the derivatization chemical reaction in-silico where the derivatization reaction is silylation. There are a number of software packages that could be used to simulate the derivatization reaction in-silico based on input of the 2D structural formula of tryptophan and of the derivatisation agent. An exemplary software package is Mass Frontier. Another exemplary software package is RDKit. The chemical reactions may be constructed using a SMARTS-based naming format, as is known in the art.

[0115] The structural formula of each of the two predicted derivatives are shown in Step 102. As shown in Step 102, those predicted derivatives are isomers having the same molecular formula but different structural formula. This is because the known precursor (tryptophan) has multiple functional groups (active hydrogens) and the derivatisation agent would react with different functional groups to form the different isomers.

[0116] In Step 102 of the method, one or more parameter(s) obtained from the structural formula of each predicted derivative can be input to the machine learning model to generate a corresponding retention parameter. Parameter(s) obtained from the structural formula may be those discussed above (e.g. atomic attributes and / or bond attributes and / or global attributes and / or the structural formula). In this exemplary implementation, the retention parameter is a retention index. As discussed above, the structural formula is input as a 2D structure as shown in Figure 4, step 103, for example, as a MOL-file. As described above in accordance with Figure 1 , the machine learning model generates a retention index based on the structural formula. The retention index generated is dimensionless and independent of measurement conditions.

[0117] The predicted derivatives are ordered according to their retention indices as shown in step 103 of Figure 4. In this case, the predicted derivatives are ordered from their lowest to highest retention indices where the lowest retention index corresponds to the shortest retention time.

[0118] As shown in Figure 4, step 104, the order of the candidate compounds (which are the predicted derivatives) according to their retention index corresponds directly to the order of the corresponding peaks according to their retention time (i.e. the order of the peaks as they appear in the chromatogram from left to right). Step 104 involves assigning peaks of the chromatogram to the candidate compounds based on this correlation between the order of the peaks according to their retention time to the order of the candidate compounds according to their retention parameter such that the candidate compounds having the retention parameter indicative of the shortest retention time is assigned to the peak with the shortest retention time. Therefore, which peak corresponds to which isomer can be readily determined without needing to perform mass spectroscopy.

[0119] Figure 5 is a second exemplary implementation of the method of the claimed invention. In Figure 5, the known precursor is cimifugin having the structural formula shown in the step 101 section of the Figure 5. This structural formula is provided as a 2D representation. The known precursor has multiple functional groups. Those functional groups are active hydrogens that are reactive with a derivatization agent to form the derivatives. In this case, the derivatization reaction is silylation.

[0120] Performing chromatography to obtain a chromatogram of the sample would be well known in the art. As shown in the step 101 section of Figure 5, the resulting chromatogram has four peaks.

[0121] The skilled person with knowledge of the known precursor (cimifugin) and knowledge of the derivatization agent (BSTFA) would be able to obtain the candidate compounds, which in this case comprise predicted derivatives of the known precursor where the derivatization is silylation.

[0122] In this example, the structural formula of each of the predicted derivatives of cimifugin is determined by simulating the derivatization chemical reaction in-silico where the derivatization reaction is silylation. As discussed above, there are a number of software packages that could be used to simulate the derivatization reaction in-silico based on input of the 2D structural formula of cimifugin and of the derivatisation agent. An exemplary software package is Mass Frontier. Another exemplary software package is RDKit. The chemical reactions may be constructed using a SMARTS-based naming format, as is known in the art.

[0123] Predicted derivatives of the known precursor obtained by silylation of the known precursor may be determined by simulating the chemical reaction of the derivatization in-silico. As discussed above in respect of Figure 4, there are a number of software packages that simulate chemical reactions in-silico based on input of the 2D structural formula of tryptophan and the derivatisation agent to the software package. An exemplary software package is Mass Frontier. Another exemplary software package is RDKit. The chemical reactions may be constructed using a SMARTS-based naming format, as is known in the art.

[0124] In this example, three predicted derivatives are produced in-silico. Those three predicted derivatives are shown in Step 102 of Figure 5. As there are four peaks in the chromatogram, the skilled person would understand that the sample is also likely to comprise the known precursor in its underivatized form and so a further candidate compound is the known precursor in its underivatized form. The structural formula of the known precursor in its derivatized form is also shown in step 102 of Figure 5.

[0125] As shown in Step 102, the predicted derivatives are isomers having the same molecular formula but different structural formula. This is because the known precursor has multiple functional groups (active hydrogens) and the derivatisation agent has reacted with different functional groups to form the different isomers.

[0126] In Step 103 of the method, one or more parameter(s) obtained from the structural formula of each candidate compound can be input to the machine learning model to generate a corresponding retention parameter, which in this exemplary implementation is a retention index. The one or more parameter(s) obtained from the structural formula may be those discussed above (e.g. atomic attributes and / or bond attributes and / or global attributes and / or the structural formula). As discussed above, the 2D structural formula is input as , for example, as a MOL-file. Certain parameters derived from or obtained from the MOL-file may be input to and used by the model to generate the retention parameter, as discussed above. As described above in accordance with Figure 1 , the machine learning model generates a retention index based on one or more parameter(s) obtained from the structural formula. The retention index generated is dimensionless and independent of measurement conditions.

[0127] The candidate compounds are ordered according to their retention indices as shown in Step 103 of Figure 5 e.g. lowest to highest retention indices. In this case, the lower retention index correlates to the shortest retention time.

[0128] As shown in Figure 5, step 104, the order of the candidate compounds according to their retention index corresponds directly to the order of the corresponding peaks according to their retention time (i.e. the order of the peaks as they appear in the chromatogram from left to right). Step 104 involves assigning peaks of the chromatogram to the candidate compounds based on this correlation between the order of the peaks according to their retention time to the order of the candidate compounds according to their retention parameter such that the candidate compounds having the retention parameter indicative of the shortest retention time is assigned to the peak with the shortest retention time.

[0129] Therefore, which peak corresponds to which isomer can be readily determined without needing to perform mass spectroscopy.

[0130] The methods described herein may be implemented as a computer program or programmable or programmed logic configured to perform the method when operated by a processor. In other words, the methods described herein may be implemented as one or more corresponding modules as hardware and / or software. For example, the methods may be implemented as one or more software components for execution by a processor of the system. Alternatively, the above-mentioned functionality may be implemented as hardware, such as on one or more field-programmable-gate-arrays (FPGAs), and / or one or more application-specific-integrated-circuits (ASICs), and / or one or more digital-signal-processors (DSPs), and / or other hardware arrangements.

[0131] It will be appreciated that, insofar as embodiments of the disclosure are implemented by a computer program, then a storage medium and a transmission medium carrying the computer program form aspects of the disclosure. The computer program may have one or more program instructions, or program code, that, when executed by a processor, causes an embodiment of the disclosure to be carried out. The term “program”, as used herein, may be a sequence of instructions designed for execution on a processor, and may include a subroutine, a function, a procedure, a module, an object method, an object implementation, an executable application, an applet, a servlet, source code, object code, a shared library, a dynamic linked library, and / or other sequences of instructions designed for execution on a computer system. The storage medium may be a magnetic disc (such as a hard drive or a floppy disc), an optical disc (such as a CD-ROM, a DVD-ROM or a BluRay disc), or a memory (such as a ROM, a RAM, EEPROM, EPROM, Flash memory or a portable / removable memory device), etc. The transmission medium may be a communications signal, a data broadcast, a communications link between two or more computers, etc.

[0132] The computer program comprising instructions that, when executed by a processor, cause the processor to perform the methods described herein may form part of a data analysis system. The data analysis system may further comprise a memory storing data, such as the chromatogram to be analysed. The data analysis system may be part of an apparatus comprising a chromatography instrument. The chromatography instrument may be a gas chromatography instrument or a liquid chromatography instrument. Such chromatography instruments are well known in the art. Typically, chromatography instruments comprise a chromatography column configured to separate the sample into its compounds, an injector configured to inject the sample into the chromatography column and a detector configured to detect the separated compounds. The data analysis system may be configured to receive the chromatogram from the chromatography instrument. As used herein, including in the claims, unless the context indicates otherwise, singular forms of the terms herein are to be construed as including the plural form and vice versa. For instance, unless the context indicates otherwise, a singular reference herein including in the claims, such as "a" or "an" (such as an analogue to digital convertor) means "one or more" (for instance, one or more analogue to digital convertor). Throughout the description and claims of this disclosure, the words "comprise", "including", "having" and "contain" and variations of the words, for example "comprising" and "comprises" or similar, mean "including but not limited to", and are not intended to (and do not) exclude other components.

[0133] The use of any and all examples, or exemplary language ("for instance", "such as", "for example" and like language) provided herein, is intended merely to better illustrate the disclosure and does not indicate a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any nonclaimed element as essential to the practice of the disclosure.

[0134] Any steps described in this specification may be performed simultaneously unless stated or the context requires otherwise.

[0135] All of the aspects and / or features disclosed in this specification may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. As described herein, there may be particular combinations of aspects that are of further benefit, for example in respect of the jet disruptor in conjunction with a sealing valve and / or an offset of the axial bore and downstream aperture. In particular, the preferred features of the disclosure are applicable to all aspects of the disclosure and may be used in any combination. Likewise, features described in non-essential combinations may be used separately (not in combination).

Claims

CLAIMS1 .A method for analysis of a chromatogram, comprising: obtaining a chromatogram of a sample comprising a mixture of compounds, wherein the mixture of compounds comprises a plurality of derivatives of a known precursor; providing one or more parameter(s) obtained from a structural formula of each of a plurality of candidate compounds to a machine learning model and generating using the machine learning model a retention parameter indicative of retention time for each of the candidate compounds as output, wherein a plurality of the candidate compounds are predicted derivatives of the known precursor; ordering the candidate compounds according to their retention parameter; and assigning peaks of the chromatogram to the candidate compounds based on correlating an order of the peaks according to their retention time to the order of the candidate compounds according to their retention parameter such that the candidate compound having the retention parameter indicative of the shortest retention time is assigned to the peak with the shortest retention time.

2. The method of claim 1 , wherein the one or more parameter(s) comprise the structural formula.

3. The method of claim 1 or claim 2, wherein the one or more parameter(s) comprise atomic attribute(s) of atom(s) of the structural formula and / or bond attributes of one or more pair(s) of atoms of the structural formula.

4. The method of claim 3, wherein the atomic attribute(s) of each atom comprise one or more of hybridisation of the atom, atom type, formal charge of the atom, hydrogen bonding type of the atom, degree of the atom, number of hydrogens the atom is connected to, whether the atom is part of an aromatic ring, chirality of the atom and / or partial charge of the atom.

5. The method of claim 3 or claim 4, wherein the bond attribute(s) of each bond comprise one or more of bond type, bond conjugation, stereo-configuration of the bond and / or if the bond is between atoms in the same ring.

6. The method of any preceding claim, wherein the mixture of compounds comprises the known precursor and wherein one of the candidate compounds is the known precursor,optionally wherein the mixture of compounds consists of the known precursor and derivatives of the known precursor.

7. The method of any one of claim 1 to 5, wherein the mixture of compounds consists of derivatives of the known precursor.

8. The method of any preceding claim, wherein the number of peaks in the chromatogram is the same as the number of candidate compounds.

9. The method of any preceding claim, wherein the retention parameter is an estimated retention index.

10. The method of any preceding claim, wherein the retention parameter is an estimated retention time.11 . The method of any preceding claim, wherein the structural formula of each of the predicted derivatives are obtained by generating derivatives of the known precursor in-silico.

12. The method of any preceding claim, wherein at least two of the plurality of derivatives of the known precursor are isomers.

13. The method of any preceding claim, wherein the structural formula of each of the predicted derivatives is obtained from a database.

14. The method of any preceding claim, wherein the method further comprises obtaining the sample by derivatising the known precursor using a derivatization agent.

15. The method of any preceding claim, wherein the known precursor comprises a plurality of functional groups.

16. The method of any preceding claim further comprising performing gas chromatography or liquid chromatography to obtain the chromatogram.

17. The method of any preceding claim further comprising calculating a regression between the retention time of the peaks of the chromatogram and the retention parameter of the corresponding assigned candidate compounds.

18. A computer program comprising instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 17.

19. An apparatus, comprising: a chromatography instrument configured to obtain the chromatogram; and a data analysis system configured to carry out the method of any one of claims 1 to 17 using the chromatogram obtained by the chromatography instrument.

Citation Information

Patent Citations

  • Method for establishing list of environmental risk aromatic compounds

    CN117612629A