Spectrometric analysis
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- WATERS TECH IRELAND LIMITED IE
- Filing Date
- 2024-06-14
- Publication Date
- 2026-04-22
AI Technical Summary
Existing spectrometric methods are inadequate for accurately characterizing samples that comprise mixtures of different components, such as healthy and cancerous tissue, as they struggle to determine the proportion of each component in a sample, especially when the mixture is in an unknown proportion.
A method of spectrometric analysis is developed that uses a probabilistic model, derived using a Bayesian approach, to determine the proportion of components in a sample from spectrometric data. This model provides an expression for the probability of the sample comprising a mixture of components in a particular proportion, allowing for the determination of component ratios or probabilities using sample spectrometric data and pre-determined statistical properties of individual components.
The method effectively determines the proportion of different tissue types or contaminants in a sample, providing accurate results even when the mixture is in an unknown proportion, improving the characterization of mixed samples in applications like intraoperative mass spectrometry and histopathology.
Smart Images

Figure EP2024066689_19122024_PF_FP_ABST
Abstract
Description
[0001] SPECTROMETRIC ANALYSIS
[0002] CROSS-REFERENCE TO RELATED APPLICATION
[0003] This application claims priority from and the benefit of United Kingdom patent application No. 2309112.7 filed on 16 June 2023, and United Kingdom patent application No. 2403714.5 filed on 14 March 2024, the entire content of which is incorporated herein by reference.
[0004] FIELD OF THE INVENTION
[0005] The present invention relates generally to spectrometry and in particular to methods of spectrometric analysis in order to characterise samples.
[0006] BACKGROUND
[0007] In known arrangements, a sample obtained from a target substance is ionised so as to produce analyte ions. The analyte ions are then subjected to mass and / or ion mobility analysis so as to produce sample spectra. The sample spectra are then subjected to spectrometric analysis in order to characterise the sample. For example, the sample spectra may be input to a spectrometric model which classifies the sample into one of plural different sample classes.
[0008] In some cases, a sample may comprise a mixture of different components. For example, in intraoperative mass spectrometry, a tissue sample may comprise healthy tissue, cancerous tissue, or a mixture of healthy tissue and cancerous tissue.
[0009] It is desired to provide improved methods of spectrometric analysis, particularly in order to characterise a sample that may comprise a mixture of different components. SUMMARY
[0010] According to an aspect of the present invention, there is provided a method of spectrometric analysis comprising: providing a model for determining, from spectrometric data representing a sample that comprises a mixture of (at least) a first component and a second component, information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample; obtaining sample spectrometric data from a sample that may comprise a mixture of (at least) the first component and the second component; and using the model and the sample spectrometric data to determine information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample.
[0011] Embodiments of the present invention relate to spectrometric analysis to characterise a sample that may comprise a mixture of (at least) a first component and a second component in an unknown proportion. For example, the sample may comprise a mixture of different tissue types, e.g. normal tissue and tumour tissue. In embodiments, sample spectrometric data obtained from the sample is used, together with the model, to determine information relating to a proportion in which the first and / or second component (e.g. tissue type) is present in the sample. For example, a percentage or ratio of tumour tissue relative to normal tissue may be determined, or a probability of tumour tissue being present in the sample in a particular proportion may be determined.
[0012] As will be discussed in more detail below, the inventors have found that it is possible to build a model that relates measured spectrometric data derived from a sample to a proportion in which first and / or second components (e.g. tissue types) are present in the sample. Using the model then makes it possible to determine, from sample spectrometric data derived from a sample that comprises a mixture of the first and second components in an unknown proportion, an indication of the proportion. This can be particularly useful, for example, where a sample may comprise a small proportion of tumour tissue, such as in the case of margin resection in intraoperative mass spectrometry, or in the case of histopathology mass spectrometry imaging, or for characterisation of contamination or adulteration e.g. of a food substance. The model may be a probabilistic model. Providing the model may comprise providing an (mathematical) expression. The expression may represent a probability distribution. The expression may be an expression for the probability that spectrometric data represents a sample that comprises a mixture of the first component and the second component in a particular proportion (from which the information relating to a proportion can be determined).
[0013] According to an aspect of the present invention, there is provided a method of spectrometric analysis comprising: providing an expression for the probability that spectrometric data represents a sample that comprises a mixture of a (at least) first component and a second component in a (at least one) particular proportion; obtaining sample spectrometric data from a sample that may comprise a mixture of (at least) the first component and the second component; and using the expression and the sample spectrometric data to determine information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample.
[0014] These aspects and embodiments can, and in embodiments do, include one or more, and in embodiments all, of the features of other aspects and embodiments described herein, as appropriate.
[0015] The expression may be derived (analytically) using a Bayesian approach. Providing the expression may comprise: providing a prior probability distribution of the proportion; providing a likelihood function of spectrometric data given the proportion; and determining a joint probability distribution of the spectrometric data and the proportion from the prior probability distribution and the likelihood function. Other Bayesian approaches may be possible.
[0016] The joint probability distribution may be determined by determining the product of the prior probability distribution and the likelihood function. The expression may be, or be based on, the joint probability distribution. The expression may be, or be based on, a log of the joint probability distribution. The expression may be, or be based on, equation (4) or (10) below.
[0017] A prior probability distribution of the proportion can be any suitable probability distribution, and may be based on any suitable prior knowledge regarding the proportion. The prior probability distribution may be experimentally determined (fitted to data) and / or determined theoretically. For example, parameter values of the prior probability distribution may be determined from (e.g. typical or historical) data, e.g. for a particular application, and / or determined theoretically. The prior probability distribution may comprise a uniform probability distribution. The prior probability distribution may comprise an exponential probability distribution. The prior probability distribution may comprise a Dirac delta distribution. Other probability distributions may be possible.
[0018] The prior probability distribution may be selected based on the (type of) analysis being performed. Accordingly, the model (e.g. expression) may be selected based on the (type of) analysis being performed, e.g. based on the (type of) sample being analysed and / or based on the (type of) sample spectrometric data obtained.
[0019] A likelihood function can be any suitable probability distribution. The likelihood function may be experimentally determined (fitted to data) and / or determined theoretically. For example, parameter values of the likelihood function may be determined from data, and / or determined theoretically. The likelihood function (and thus model / expression) may be based on statistical properties of component spectrometric data representing individual components of the mixture. The likelihood function (and thus model / expression) may be based on component spectrometric data following particular probability distributions. For example, component spectrometric data may (be assumed to) follow Gaussian statistics, and the likelihood function may be based on Gaussian statistics. The likelihood function may be, or be based on, equation (3) below. Other probability distributions, such as Cauchy, may be possible.
[0020] Component statistical properties (e.g. parameter values) may be determined in a same or different experiment to an experiment in which the information relating to a proportion is determined. For example, component statistical properties may be predetermined (e.g. as part of a calibration procedure) and stored in storage for later use. The method may comprise retrieving information indicative of component statistical properties from storage, and using the component statistical properties to determine the information relating to a proportion. Similarly, providing the model (e.g. expression) may comprise retrieving information and / or instructions defining the model (e.g. expression) from (the) storage.
[0021] First information indicative of statistical properties of component spectrometric data representing samples comprising (only) the first component may be provided, e.g. retrieved from storage. Second information indicative of statistical properties of component spectrometric data representing samples comprising (only) the second component may be provided, e.g. retrieved from storage. The expression, the sample spectrometric data, the first information and the second information may then be used to determine the information relating to a proportion.
[0022] The statistical properties (e.g. parameter values) may be obtained by fitting probability distributions to reference spectrometric data obtained from reference samples of the individual components (e.g. in calibration experiments). Providing the first information may comprise: obtaining first reference spectrometric data from first reference samples comprising (only) the first component; and fitting one or more first probability distributions to the first reference spectrometric data. Providing the second information may comprise: obtaining second reference spectrometric data from second reference samples comprising (only) the second component; and fitting one or more second probability distributions to the second reference spectrometric data.
[0023] The fitting may comprise extracting a set of features from spectrometric data, and fitting a respective probability distribution for each feature. A feature may be a sample bin, a peak, or a group of peaks.
[0024] The fitting may comprise binning (first and / or second) reference spectrometric data into plural sample bins, and then fitting a respective probability distribution to data (intensity values) binned into each sample bin. The sample bins may comprise mass to charge ratio or ion mobility (e.g. drift time) bins.
[0025] The fitting may comprise peak detecting (first and / or second) reference spectrometric data; obtaining, retrieving, or determining a target peak list; associating (a subset of) the detected peaks with the target peaks; and fitting a respective probability distribution to data (intensity values) assigned to each target peak. The peak detection may be performed in the mass-to-charge ratio domain (mass spectral peak detection), in the ion mobility or drift time domain (ion mobility spectral peak detection), or in both domains (e.g. in an ion mobility-mass spectrometry experiment).
[0026] The fitting may comprise fitting Gaussian distributions, e.g. by determining mean and standard deviation (or variance) values. A mean and standard deviation (or variance) value may be determined for each feature, e.g. sample bin or peak. The (stored) statistical properties (e.g. parameter values) may comprise the mean and standard deviation (or variance) values. Other probability distributions may be fitted. The (stored) statistical properties may (further) comprise correlation information. The fitting may comprise calculating correlation information (e.g. covariances) between features, e.g. bins or peaks.
[0027] The model (e.g. expression) and sample spectrometric data may be used to determine the information relating to a proportion in any suitable manner. The expression may be evaluated using the sample spectrometric data, and optionally using the (stored) statistical properties, e.g. first information and / or the second information. A probability distribution (that the expression represents) may be sampled, e.g. using a MCMC (Markov chain Monte Carlo) method. A value of the proportion may be determined as the mean of the samples, and optionally an uncertainty on the value of the proportion may be determined from the variance of the samples. Alternatively, a value of the proportion may be determined as the maximum of the probability distribution (e.g. determined analytically or numerically), and optionally an uncertainty on a value of the proportion may be determined directly from the probability distribution. Other methods may be possible.
[0028] The model (e.g. expression) may be for determining, from spectrometric data representing a sample that comprises a mixture of more than two components, information relating to a proportion in which any one or more of the components are present in the sample. For example, the model (e.g. expression) may relate to a mixture of three, four, or more components. Similarly, information relating to more than one proportion may be determined.
[0029] A proportion may be a proportion of any one or more components relative to any one or more (e.g. all) components. For example, a (the) proportion may be a proportion of the first component relative to the second component, a proportion of the second component relative to the first component, a proportion of the first and / or second component relative to one or more other components, a proportion of the first and / or second and / or other component relative to all components, etc..
[0030] The information relating to a proportion may indicate one or more of: (i) a value of the proportion; (ii) an uncertainty on a value of the proportion; (iii) a confidence interval on a value of the proportion; (iv) a probability of the proportion being less than, and / or equal to, and / or greater than, a threshold value; (v) a probability of the proportion being within a range of values; (vi) a probability of (at least) the first component and / or the second component being present in the sample in a proportion less than, and / or equal to, and / or greater than, a threshold value; (vii) a probability of (at least) the first component and / or the second component being present in the sample in a proportion within a range of values; and (viii) a class of plural classes representing different values of the proportion. A value of the proportion may, e.g., be a ratio or percentage (that a component is determined to be present in the sample (relative to another component)).
[0031] In embodiments, the model is a machine learning model, such as an artificial neural network. The machine learning model may be trained to be able to characterise a sample that comprises a mixture of the first component and the second component, e.g. to be able to determine the information relating to a proportion.
[0032] According to an aspect, there is provided a method of spectrometric analysis comprising: providing a machine learning model for determining, from spectrometric data representing a sample that comprises a mixture of (at least) a first component and a second component, information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample; obtaining sample spectrometric data from a sample that may comprise a mixture of (at least) the first component and the second component; and using the machine learning model and the sample spectrometric data to determine information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample.
[0033] According to an aspect, there is provided a method of training a machine learning model to be able to characterise a sample that may comprise a mixture of (at least) a first component and a second component using spectrometric data representing the sample, the method comprising: obtaining first reference spectrometric data from first reference samples comprising the first component; obtaining second reference spectrometric data from second reference samples comprising the second component; fitting one or more first probability distributions to the first reference spectrometric data; fitting one or more second probability distributions to the second reference spectrometric data; using the one or more first probability distributions and the one or more second probability distributions to generate training data representative of spectrometric data derived from a sample comprising a mixture of (at least) the first component and the second component; and using the generated training data to train a machine learning model to be able to characterise a sample that may comprise a mixture of (at least) the first component and the second component using spectrometric data representing the sample.
[0034] According to an aspect, there is provided a method of characterising a sample comprising: obtaining sample spectrometric data from a sample; and using the sample spectrometric data and a machine learning model trained as described above to characterise the sample.
[0035] These aspects and embodiments can, and in embodiments do, include one or more, and in embodiments all, of the features of other aspects and embodiments described herein, as appropriate. For example, the machine learning model may be trained to be able to determine the information relating to a proportion as described above, the fitting may be as described above, etc..
[0036] These aspects and embodiments allow the generation of “synthetic” training spectrometric data to train a machine learning model to be able to characterise mixed samples. This may be particularly useful where it is difficult to obtain sufficient training data to train a machine learning model to be able to characterise mixed samples, which can be the case where there is a limited availability of mixed reference samples, and / or where annotation of mixed reference samples is subjective and unreliable.
[0037] The machine learning model can be any suitable machine learning model. The machine leaning model may be an artificial neural network. A neural network may take any desired and suitable form. A neural network may comprise one or more fully connected networks, one or more activation layers, etc.. A neural network may comprise an input layer that receives an input, one or more intermediate or “hidden” layers, and an output layer that provides a result. The neural network may comprise a convolutional neural network (CNN), for example having one or more convolutional layers (e.g. that each apply one or more convolution operations to generate an output for the layer), and / or one or more pooling layers (e.g. that each pool or aggregate sets of input values to generate an output from the layer), and / or one or more deconvolution layers, etc.. According to an aspect, there is provided a method of providing a model for characterising a sample that may comprise a mixture of (at least) a first component and a second component using spectrometric data representing the sample, the method comprising: obtaining first reference spectrometric data from first reference samples comprising the first component; obtaining second reference spectrometric data from second reference samples comprising the second component; fitting one or more first probability distributions to the first reference spectrometric data; fitting one or more second probability distributions to the second reference spectrometric data; using the one or more first probability distributions and the one or more second probability distributions to synthesise data representative of spectrometric data derived from a sample comprising a mixture of (at least) the first component and the second component; and using the synthesised data to build and / or test a model for characterising a sample that may comprise a mixture of (at least) the first component and the second component using spectrometric data representing the sample.
[0038] According to an aspect, there is provided a method of spectrometric analysis comprising characterising a sample using a model provided as described above.
[0039] The synthesised / training data may be generated by sampling the one or more first probability distributions and the one or more second probability distributions, e.g. using a MCMC (Markov chain Monte Carlo) method. Other methods may be possible.
[0040] The synthesised / training data may comprise one or more synthesised / training spectra (each) having plural features, e.g. sample bins or peaks. The one or more first probability distributions may comprise a first probability distribution corresponding to each feature, e.g. sample bin or peak. The one or more second probability distributions may comprise a second probability distribution corresponding to each feature, e.g. sample bin or peak. The first and / or second probability distributions may comprise correlations between features.
[0041] A (each) synthesised / training spectrum may be generated by, for each feature (e.g. sample bin or peak) of the synthesised / training spectrum: sampling the first probability distribution corresponding to the respective feature to determine a first (“synthetic”) intensity value; sampling the second probability distribution corresponding to the respective feature to determine a second (“synthetic”) intensity value; and summing the first intensity value and the second intensity value to determine an (“synthetic”) intensity value for the respective feature. The first intensity value and the second intensity value may be summed according to a weighted sum, where the weight may be selected to represent a particular proportion in which the first and / or second component is present.
[0042] Synthesised / training data may be generated using correlations, e.g. between different features.
[0043] In the case of more than two components, further reference spectrometric data from one or more further reference samples comprising one or more further components may be obtained, and one or more further probability distributions may be fitted to the further reference spectrometric data and used (e.g. sampled) to synthesise the data, etc..
[0044] Spectrometric data may be obtained from a sample in any suitable manner. Sample spectrometric data and reference spectrometric data may be obtained in substantially the same manner. Spectrometric data may be obtained by analysing ions derived from a sample using an analyser, e.g. by mass spectrometry and / or ion mobility spectrometry. The (sample and / or reference) spectrometric data may comprise mass spectrometric data and / or ion mobility spectrometric data, e.g. representing mass to charge ratio, time of flight, ion mobility, differential ion mobility, drift time, collision cross section (CCS), etc..
[0045] Obtaining spectrometric data from a sample may comprise obtaining the sample using a sampling device. The sample may be obtained from a target. The sample may be obtained from one or more regions of a target. Obtaining spectrometric data from a sample may comprise generating a plurality of analyte ions from the sample using an ion source.
[0046] The sampling device may comprise or form part of an ion source. The sampling device may comprise or form part of an ambient ionisation or ambient ion source. The sampling device may comprise a surgical device.
[0047] The sample may comprise an aerosol, smoke or vapour sample. Obtaining spectrometric data from a sample may comprise generating the aerosol, smoke or vapour sample using a sampling device. Generating the aerosol, smoke or vapour sample may comprise contacting a target with one or more electrodes. Generating the aerosol, smoke or vapour sample may comprise applying an AC or RF voltage to the one or more electrodes in order to generate the aerosol, smoke or vapour sample. Generating the aerosol, smoke or vapour sample may comprise direct evaporation or vaporisation of target material from a target by Joule heating or diathermy. Generating the aerosol, smoke or vapour sample may comprise irradiating a target with a laser.
[0048] Spectrometric data may be obtained using an ambient ionisation technique, such as rapid evaporative ionisation mass spectrometry (“REIMS”), desorption electrospray ionization (“DESI”), or matrix-assisted laser desorption / ionization (“MALDI”).
[0049] The target may comprise inorganic matter and / or non-biological matter. The target may be from or form part of a human or non-human animal subject (e.g., a patient). The target may comprise organic matter, biological tissue, biological matter, a bacterial colony or a fungal colony. The biological tissue may comprise human tissue or non-human animal tissue. The biological tissue may comprise in vivo biological tissue. The biological tissue may comprise ex vivo biological tissue. The biological tissue may comprise in vitro biological tissue.
[0050] The first component may be a first (biological) tissue type. The second component may be a second, different (biological) tissue type. The first tissue type may be normal (e.g. healthy) tissue, and the second tissue type may be abnormal (e.g. tumour) tissue.
[0051] The first and / or second tissue type may be selected from: (i) adrenal gland tissue, appendix tissue, bladder tissue, bone, bowel tissue, brain tissue, breast tissue, bronchi, coronal tissue, ear tissue, esophagus tissue, eye tissue, gall bladder tissue, genital tissue, heart tissue, hypothalamus tissue, kidney tissue, large intestine tissue, intestinal tissue, larynx tissue, liver tissue, lung tissue, lymph nodes, mouth tissue, nose tissue, pancreatic tissue, parathyroid gland tissue, pituitary gland tissue, prostate tissue, rectal tissue, salivary gland tissue, skeletal muscle tissue, skin tissue, small intestine tissue, spinal cord, spleen tissue, stomach tissue, thymus gland tissue, trachea tissue, thyroid tissue, ureter tissue, urethra tissue, soft and connective tissue, peritoneal tissue, blood vessel tissue and / or fat tissue; (ii) grade I, grade II, grade III or grade IV cancerous tissue; (iii) metastatic cancerous tissue; (iv) mixed grade cancerous tissue; (v) a sub-grade cancerous tissue; (vi) healthy or normal tissue; and / or (vii) cancerous or abnormal tissue. Other tissue types may be possible.
[0052] The first component may be a substance such as food, and the second component may be a contaminant or adulterant.
[0053] The model may be a model for determining, from spectrometric data representing a sample that comprises a mixture of more than two components, information relating to a proportion in which any one or more of the components are present in the sample. For example, the model may relate to a mixture of three, four, or more components.
[0054] The method may be computer implemented. The method may be non- surgical and / or non-therapeutic and / or non-diagnostic. The method may be a method of mass spectrometry and / or ion mobility spectrometry.
[0055] According to an aspect, there is provided a method of operating an analytical instrument, comprising a method as described above. The analytical instrument may be a mass spectrometer and / or an ion mobility spectrometer.
[0056] According to an aspect, there is provided a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out a method as described above.
[0057] According to an aspect, there is provided a (non-transitory) computer- readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out a method as described above.
[0058] According to an aspect, there is provided a spectrometric analysis system comprising control circuitry arranged and adapted to perform a method as described above.
[0059] According to an aspect, there is provided an analytical instrument comprising control circuitry arranged and adapted to perform a method as described above.
[0060] According to an aspect, there is provided a mass spectrometer and / or ion mobility spectrometer comprising control circuitry arranged and adapted to perform a method as described above.
[0061] Thus, according to aspects, there is provided a spectrometric analysis system or an analytical instrument, such as a mass spectrometer and / or ion mobility spectrometer, comprising: control circuitry arranged and adapted to: provide a model for determining, from spectrometric data representing a sample that comprises a mixture of (at least) a first component and a second component, information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample; obtain sample spectrometric data from a sample that may comprise a mixture of (at least) the first component and the second component; and use the model and the sample spectrometric data to determine information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample.
[0062] According to aspects, there is provided a spectrometric analysis system or an analytical instrument, such as a mass spectrometer and / or ion mobility spectrometer, comprising: control circuitry arranged and adapted to: provide an expression for the probability that spectrometric data represents a sample that comprises a mixture of (at least) a first component and a second component in a (at least one) particular proportion; obtain sample spectrometric data from a sample that may comprise a mixture of (at least) the first component and the second component; and use the expression and the sample spectrometric data to determine information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample.
[0063] According to aspects, there is provided a spectrometric analysis system or an analytical instrument, such as a mass spectrometer and / or ion mobility spectrometer, comprising: control circuitry arranged and adapted to: provide a machine learning model for determining, from spectrometric data representing a sample that comprises a mixture of (at least) a first component and a second component, information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample; obtain sample spectrometric data from a sample that may comprise a mixture of (at least) the first component and the second component; and use the machine learning model and the sample spectrometric data to determine information relating to a (at least one) proportion in which (at least) the first component and / or the second component is present in the sample. According to aspects, there is provided a spectrometric analysis system or an analytical instrument, such as a mass spectrometer and / or ion mobility spectrometer, comprising: control circuitry arranged and adapted to: obtain first reference spectrometric data from first reference samples comprising the first component; obtain second reference spectrometric data from second reference samples comprising the second component; fit one or more first probability distributions to the first reference spectrometric data; fit one or more second probability distributions to the second reference spectrometric data; use the one or more first probability distributions and the one or more second probability distributions to synthesise data representative of spectrometric data derived from a sample comprising a mixture of (at least) the first component and the second component; and use the synthesised data to build and / or test a model for characterising a sample that may comprise a mixture of (at least) the first component and the second component using spectrometric data representing the sample.
[0064] According to aspects, there is provided a spectrometric analysis system or an analytical instrument, such as a mass spectrometer and / or ion mobility spectrometer, comprising: control circuitry arranged and adapted to: obtain first reference spectrometric data from first reference samples comprising the first component; obtain second reference spectrometric data from second reference samples comprising the second component; fit one or more first probability distributions to the first reference spectrometric data; fit one or more second probability distributions to the second reference spectrometric data; use the one or more first probability distributions and the one or more second probability distributions to generate training data representative of spectrometric data derived from a sample comprising a mixture of (at least) the first component and the second component; and use the generated training data to train a machine learning model to be able to characterise a sample that may comprise a mixture of (at least) the first component and the second component using spectrometric data representing the sample.
[0065] The control circuitry may be configured to control a sampling device and / or ion source and / or analyser of the system or instrument, e.g. as described above. The control circuitry may comprise, or be in communication with, storage (e.g. memory) storing instructions and / or information for carrying out the method. For example, information and / or instructions defining the model (e.g. expression) and / or the first and / or second information may be stored in the storage (and retrieved therefrom in order to determine the information relating to a proportion).
[0066] Each aspect can, and in embodiments does, include one or more, and in embodiments all, of the features of other aspects and embodiments described herein, as appropriate.
[0067] BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Various embodiments of the present invention will now be described by way of example only and with reference to the accompanying drawings, in which:
[0069] Figure 1 shows an analytical instrument that may be operated in accordance with embodiments;
[0070] Figure 2 shows an ambient ionisation mass spectrometry system that may be operated in accordance with embodiments;
[0071] Figure 3 shows a spectrometric analysis method in accordance with embodiments;
[0072] Figure 4A and Figure 4B show reference spectrometric data obtained from samples of individual components in accordance with embodiments;
[0073] Figure 5A and Figure 5B show Gaussian probability distributions fitted to reference spectrometric data in accordance with embodiments;
[0074] Figures 6A, 6B, 60 and 6D show results of spectrometric analyses in accordance with embodiments;
[0075] Figure 7A shows a method of training a machine learning model in accordance with embodiments, and Figure 7B shows a spectrometric analysis method that uses a machine learning model trained according to the method of Figure 7A; Figure 8 shows a neural network in accordance with embodiments; and Figures 9A, 9B and 9C show results of spectrometric analyses in accordance with embodiments.
[0076] DETAILED DESCRIPTION
[0077] Figure 1 shows schematically an analytical instrument which may be operated in accordance with various embodiments. As shown in Figure 1, the analytical instrument comprises an ion source 10, one or more functional components 20 that are arranged downstream from the ion source 10, and an analyser 30 that is arranged downstream from the ion source 10 and from the one or more functional components 20.
[0078] It should be noted that Figure 1 is merely schematic, and that the analytical instrument may (and in various embodiments does) include other components, devices and functional elements to those shown in Figure 1.
[0079] The ion source 10 may be configured to generate ions, for example by ionising an analyte. The ion source 10 may comprise any suitable ion source. The analytical instrument may optionally comprise a chromatography or other separation device (not shown in Figure 1) upstream of (and coupled to) the ion source 10. The analytical instrument may optionally comprise a sampling device (not shown in Figure 1) upstream of (and coupled to) the ion source 10 and configured to provide a sample to the ion source 10 for ionisation.
[0080] The analyser 30 may be configured to analyse ions, so as to determine (measure) one or more of their physico chemical properties, such as their mass to charge ratio, time of flight, (ion mobility) drift time and / or collision cross section (CCS), differential ion mobility. The analyser 30 may comprise a mass analyser (that is configured to determine the mass to charge ratio or time of flight of ions) and / or an ion mobility analyser (that is configured to determine e.g. the ion mobility drift time or collision cross section (CCS) or (differential) ion mobility of ions). The mass analyser may, for example, comprise a quadrupole mass analyser or a Time of Flight mass analyser.
[0081] As shown in Figure 1, the analytical instrument may comprise a control system 40, that may be configured to control the operation of the analytical instrument, for example in the manner of the various embodiments described herein. The control system may comprise suitable control circuitry that is configured to cause the instrument to operate in the manner of the various embodiments described herein. The control system may comprise suitable processing circuitry configured to perform any one or more or all of the necessary processing and / or post-processing operations in respect of the various embodiments described herein. In various embodiments, the control system may comprise a suitable computing device, a microprocessor system, a programmable FPGA (field programmable gate array), and the like. In various embodiments, the control system comprises storage (e.g. a memory) for storing information and instructions for performing methods described herein.
[0082] As illustrated by Figure 1 the analytical instrument may be configured such that ions can be provided by (sent from) the ion source 10 to the analyser 30 via the one or more functional components 20. The one or more functional components 20 may comprise any suitable such components, devices and functional elements of an analytical instrument (mass and / or ion mobility spectrometer).
[0083] In various embodiments, the one or more functional components 20 comprise one or more ion guides and / or one or more ion traps. In various embodiments, the one or more functional components 20 comprise a mass filter, which may be configured to filter ions according to their mass to charge ratio. In various embodiments, the one or more functional components 20 comprise an activation, collision, fragmentation or reaction device configured to activate, fragment or react ions. In various embodiments, the one or more functional components 20 comprise an ion mobility separator configured to separate ions according to their ion mobility. Other functional components 20 would be possible.
[0084] Figure 2 shows an embodiment in which a sampling device is used to obtain a sample from a target, and to provide the sample to the ion source 10. Figure 2 shows a rapid evaporative ionisation mass spectrometry (“REIMS”) system in which an RF voltage is applied to one or more electrodes in order to generate an aerosol or plume of surgical smoke by Joule heating. As shown in Figure 2, bipolar forceps 1 may be brought into contact with in vivo tissue 2 of a patient 3. In the example shown in Figure 2, the bipolar forceps 1 may be brought into contact with brain tissue 2 of a patient 3 during the course of a surgical operation on the patient’s brain. Other tissues and targets are possible.
[0085] An RF voltage from an RF voltage generator 4 may be applied to the bipolar forceps 1 which causes localised Joule or diathermy heating of the tissue 2. As a result, an aerosol or surgical plume 5 is generated. The aerosol or surgical plume 5 may then be captured or otherwise aspirated through an irrigation port of the bipolar forceps 1. The aerosol or surgical plume 5 may then be passed from the irrigation (aspiration) port of the bipolar forceps 1 to tubing 6 (e.g. 1 / 8” or 3.2 mm diameter Teflon (RTM) tubing). The tubing 6 is arranged to transfer the aerosol or surgical plume 5 to an atmospheric pressure interface 7 of the analytical instrument 8.
[0086] A matrix comprising an organic solvent such as isopropanol (I PA) may be added to the aerosol or surgical plume 5 at the atmospheric pressure interface 7. The mixture of aerosol and organic solvent may then be arranged to impact upon a collision surface within a vacuum chamber of the mass spectrometer and / or ion mobility spectrometer 8. The collision surface may be heated. The aerosol may be caused to ionize upon impacting the collision surface resulting in the generation of analyte ions. The ionization efficiency of generating the analyte ions may be improved by the addition of the organic solvent. However, the addition of an organic solvent is not essential.
[0087] Analyte ions which are generated by causing the aerosol, smoke or vapour 5 to impact upon the collision surface may then be passed through the subsequent stages 20 of the analytical instrument 8, optionally fragmented, and subjected to mass analysis and / or ion mobility analysis by the analyser 30. Resulting spectrometric data may then by subjected to spectrometric analysis by the control system 40 in order to characterise (e.g. determine one or more properties of) the target 2, for example in real time. For example, it may be determined whether the target 2 comprises cancerous tissue or normal tissue.
[0088] Other arrangements are possible. For example, other sampling, ionisation, and analyser arrangements may be used. For example, other e.g. ambient ionisation techniques, such as desorption electrospray ionization (“DESI”) and matrix-assisted laser desorption / ionization (“MALDI”), can be used.
[0089] In the system described above, a sample 5 obtained from a target 2 is ionised by ion source 10 so as to generate analyte ions. The resulting analyte ions (or fragment or product ions derived from the analyte ions) are then mass and / or ion mobility analysed by the analyser 30, and the resulting mass and / or ion mobility spectrometric data may then be subjected to spectrometric analysis (by the control system 40) in order to characterise the sample / target.
[0090] Typically, the spectrometric data is input to a spectrometric model which classifies the sample into one of plural different possible sample classes, e.g. cancerous tissue or normal tissue. Such spectrometric models may typically be built using reference spectrometric data obtained e.g. by analysing pure standards, or derived from theory or exemplar spectra.
[0091] The inventors have recognised, however, that it can be the case that a sample comprises a mixture of different known components. For example, in the system described above, a sample 5 may comprise a mixture of healthy tissue and cancerous tissue. Such a mixed sample may, for example, be obtained at a boundary between cancerous and normal tissue, e.g. during resection of a tumour with a margin. Mixed samples may also be obtained in other analysis settings, such as at tissue boundaries in MS imaging experiments, such as DESI or MALDI imaging, e.g. for histopathology. A mixed sample may also be obtained in the case of contamination or adulteration of a substance, e.g. food, such as a beef burger adulterated with offal.
[0092] The inventors have found that known spectrometric models may be inadequate for characterising such mixed samples. For example, it may be desired to determine a ratio or proportion in which different components are present in a mixed sample, e.g. relative proportions of cancerous tissue and normal tissue, or a relative proportion of a contaminant or adulterant.
[0093] Various embodiments accordingly provide spectrometric models that are able to characterise a sample that comprises a mixture of different components, e.g. a mixture of different tissue types. Particular embodiments provide spectrometric models that are able to determine an indication of a proportion in which different components (e.g. tissue types) are (likely to be) present in a sample.
[0094] Figure 3 illustrates a method of characterising a sample in accordance with various embodiments. In these embodiments, a Bayesian approach is used to calculate an (mathematical) expression which can be used to determine, from spectrometric data derived from a sample that comprises a mixture of components (e.g. tissue types), and from known statistical properties of the individual components (e.g. tissue types), a proportion in which the different components (e.g. tissue types) are (likely to be) present in the sample.
[0095] As shown in Figure 3, at step 301, an (mathematical) expression for the probability of spectrometric data being derived from a mixture of different components in a particular proportion is provided. In embodiments, the expression is derived analytically using a Bayesian approach, e.g. as described below. At step 302, statistical properties of spectrometric data derived from the individual components are provided. In embodiments, the spectrometric statistical properties are derived from samples of each component individually (when not in a mixture), e.g. as described below.
[0096] At step 303, sample spectrometric data is obtained from a sample that comprises a mixture of the different components in an unknown proportion, e.g. by analysing ions derived from the sample. Then, at step 304, the expression, the component statistical information and the sample spectrometric data, are used to determine a proportion in which the different components are likely to be present in the sample.
[0097] The expression can be provided in any suitable manner. In embodiments, measured spectrometric data d, is considered to be an array of N real numbers (1 < / < / V) produced by a linear combination of the spectra of two known components f and g. In particular, the measured spectrum d, is considered to be a weighted sum of components f and g with weights a and (1 - a), respectively, where 0 < a < 1.
[0098] In the present embodiments, it is assumed that the component spectra f and g have associated (independent) Gaussian uncertainties ar and og. In particular, it is assumed that for each datapoint (bin or feature) (each / ), f and g each follow an (independent) Gaussian probability distribution, represented by means f, and g, and standard deviations On and ogrespectively. It is assumed that
[0099] In the present embodiments, it is furthermore assumed that the data can be measured perfectly. The approach according to embodiments may thus be particularly suitable for use where a mismatch between measured spectrometric data and model arises predominantly from variability in the underlying components, rather than limited accuracy of the data. In particular, the approach is suitable for detection of mixtures of tissue types where spectra arising from particular (pure) tissue types may differ from each other as a result of real differences in the samples originating from a single tissue type (e.g. biological variability or a degree of unavoidable contamination). That is, the approach according to embodiments is particularly suitable where there is a natural statistical variation (e.g. Gaussian distribution) in (each bin of) spectra obtained from each component (e.g. tissue type).
[0100] Since d is a weighted sum of the component spectra f and g, a distribution of d will be a convolution of Gaussian distributions of the individual components, and the convolution of two Gaussian distributions is another Gaussian distribution. Because means and variances add under convolution, d can thus be represented (symbolically) as: di = ,i(a) ± (Ti a)' , where are mean and standard deviation, respectively, of Gaussian distribution(s) for datapoints of d.
[0101] The weight a may be assigned any suitable prior probability distribution. In the present embodiment, a uniform prior is assigned to the weight a over the range [0, 1].
[0102] The likelihood of the data d given the parameter a is the product of contributions from each individual datapoint:
[0103] In the case of a uniform prior over a, the joint probability is equal to the likelihood. To address the issue of small numbers in numerical implementations it can be helpful to work with the logarithm of the distribution. The logarithm of the joint probability is thus:
[0104] Equation (4) (with equation (2)) can be used to determine, from spectrometric data d derived from a sample that comprises a mixture of components, and from known statistical properties of the individual components (fi, g On, Ogi), a proportion a in which the different components are (likely to be) present in the sample. The (known) statistical properties of the individual components can be provided in any suitable manner. The statistical properties may be predetermined from samples of individual components, e.g. as part of a calibration procedure, and stored for later use.
[0105] Figures 4 and 5 illustrate an example in which reference mass spectra were derived from samples comprising only normal tissue and samples comprising only tumour tissue using a rapid evaporative ionisation mass spectrometry (“REIMS”) system as described above. Figure 4A shows mass spectra derived from samples comprising only normal breast tissue, and Figure 4B illustrates mass spectra derived from samples comprising only breast tumour tissue.
[0106] The mass spectra were sum normalised, lockmass calibrated, and binned into / V bins of width 0.1 Da. Other pre-processing steps are possible. Figure 5A shows the resulting pre-processed normal tissue spectrometric data, and Figure 5B shows the resulting pre-processed tumour tissue spectrometric data. For each resulting bin, a mean of the intensities corresponding to the bin and a standard deviation of the intensities corresponding to the bin were determined. The error bars shown in Figure 5 represent the calculated mean and standard deviation values.
[0107] The calculated mean and standard deviation values for the tumour tissue were stored for use as the mean f, and standard deviation On values for component f, and the calculated mean and standard deviation values for the normal tissue were stored for use as the mean g, and standard deviation ag, values for component g.
[0108] Mass spectra from samples comprising known proportions of breast tumour tissue and normal breast tissue were obtained, and used to test the model. The sample mass spectra were pre-processed in a corresponding manner to the reference mass spectra, i.e. sum normalised, lockmass calibrated, and binned into N bins of width 0.1 Da, and the resulting pre-processed sample spectrometric data used as d, in equation (4). A MCMC (Markov chain Monte Carlo) method was then used to sample equation (4), and results are illustrated in Figure 6.
[0109] Figure 6A illustrates the results of applying the above method to a sample comprising 5.5% breast tumour and 94.5% normal breast tissue. Figure 6A shows (in the darker colour), a histogram of values of a determined by sampling equation (4) using the statistical information (means and standard deviation values, fi, g On, Ogi) illustrated in Figure 5 and the mass spectrometric data derived from the “5.5% sample”. The first moment (mean) of the resulting distribution was determined to be 5.5%, indicating that the model was able to accurately determine a value of a.
[0110] Although the mean was used to infer a value of a, other estimates, e.g. median or mode, may be determined. Furthermore, a measure of uncertainty may be determined, e.g. by determining the second moment (variance), standard deviation, or confidence interval, etc.. It would also be possible to determine other statistics from the distribution, such as other moments, or a measure of the probability of the proportion a being above or below a particular threshold.
[0111] For comparison, the method was applied to a sample comprising only normal breast tissue, and Figure 6A also shows (in the lighter colour) a histogram of values of a determined by sampling equation (4) using the statistical information (means and standard deviation values, fi, g On, ogi) illustrated in Figure 5 and mass spectrometric data derived from the sample of normal breast tissue. As can be seen in Figure 6A, good separation between the normal and tumour containing data is observed.
[0112] Figures 6B, 6C and 6D illustrate corresponding results for samples comprising 4%, 3.5% and 2% breast tumour, respectively (in the darker colour), and for comparison, results for a sample comprising only normal breast tissue (in the lighter colour). Again, the method was able to satisfactorily determine an estimated value of a.
[0113] It will be appreciated that other (numerical) methods of solving equation (4) would be possible. For example, it would be possible to evaluate (e.g. plot) Pr(d, a) for different values of a, e.g. and numerically search for the value of a that maximises Pr(d, a) (the mode or maximum a posteriori value).
[0114] Although in the spectrometric model described above, components f and g are assumed to follow Gaussian statistics, it would be possible to assume other probability distributions and / or fit other probability distributions to the reference data. For example, a generalised Cauchy distribution may be used, e.g. having the form: where 1 < C < °°. This generalised Cauchy distribution reduces to a standard Cauchy distribution for C = 1, and becomes a Gaussian distribution as C goes to infinity. The parameter D, controls the width of the distribution (in the Gaussian limit Di is simply the standard deviation) while the global value C controls the size of the tails. The value of C may be fixed in advance, or determined from the data. Optionally, the value of C could be allowed to vary from bin to bin or feature to feature.
[0115] Similarly, although in the spectrometric model described above, a is assumed to follow a uniform prior probability distribution, other distributions could be assumed and / or fitted to data. For example, a probability distribution reflecting an expected distribution of a in a given setting, e.g. resection surgery or tissue imaging in histopathology, could be used. The prior could be a continuous function of a, but might also contain a Dirac delta function at a = 0 if a significant proportion of spectra are expected to be completely normal, e.g.: where h is the proportion of spectra that are expected to be completely normal, and A is the expected proportion of “tumour” in non-normal tissue.
[0116] Although in the spectrometric model described above, each datapoint / is assumed to follow an independent probability distribution, it would be possible to assume, and / or determine from data, correlations between different datapoints.
[0117] Furthermore, it is possible to extend the model to more than two different components. In this case, the ratio of the abundance of a target component to the sum of the abundances of all components (including the target component) could be inferred, or the ratio of any two individual components of interest could be inferred.
[0118] For example, data d, can be considered to be an array of N real numbers produced by linear combination of J known components that have shapes (spectra) fji, where 1 < / < N and 1 <j< J, and for each component j, It may be assumed that 52' j = 1, measurec| perfectly, that the components f7have associated (independent) Gaussian uncertainties, OJI, and that the data are a ■ = 1 produced by a mixture of components f7with weights a,, whereJ In this case, d can be represented (symbolically) as: di = Ai (a) ± cTj(a)
[0119] (7) where:
[0120] Assigning a uniform prior to the weights a over the simplex a-11 then provides a likelihood: and logarithm of joint probability:
[0121] Alternatively, relevant knowledge concerning the prevalence of the component types could be encoded in the prior probability distribution for the weights a. For example, a non-zero probability could be assigned to the possibility a7= 0 for one or more of the components j.
[0122] Equation (10) is a multidimensional probability distribution over the simplex . This can be sampled using a Markov Chain Monte Carlo (MCMC) algorithm, and statistics concerning the relative sizes of combinations of the individual a7can be accumulated. For example, if one component k (1 < k < J) is considered to be exceptional (e.g. tumour in tissue analysis), and all remaining components are considered to be normal (e.g. healthy tissue), then accumulating the average of sampled values of dk can yield the inferred proportion of component k in the data, along with other statistics such as the standard deviation of dk, associated quantiles or confidence intervals, and (given an appropriate prior) Pr(a / < > 0). Alternatively, if there is more than one exceptional component, then the sum of the relevant a7may be accumulated. Statistics regarding any other functions of the dk that are of interest may be accumulated, for example the ratio of the abundances of two components anlam, etc..
[0123] Although in the spectrometric model described above, each datapoint / corresponds to a m / z bin, other features may be used, such as peaks, groups of peaks, etc.. For example, a feature may correspond to a group of isotopes or related compounds, or a peak resulting from a deisotoping algorithm.
[0124] In embodiments, the size of the model space may be reduced by “feature selection”, e.g. bin selection. For example, features may be ranked by crossentropy (or relative entropy) of the distributions in each bin between classes. This can allow an objective selection of a top N features that are maximally different (in a specific sense) between classes.
[0125] Various further embodiments will now be described.
[0126] As discussed above, the inventors have found that known spectrometric models may be inadequate for characterising mixed samples. For example, the inventors have found that, in the case of machine learning based spectrometric models, it can be difficult to obtain sufficient training data to train a machine learning model to be able to characterise mixed samples. For example, there can be a limited availability of mixed samples, and annotation of mixed samples can be subjective and unreliable.
[0127] Figure 7 illustrates a method of characterising a sample in accordance with various embodiments. In these embodiments, a machine learning spectrometric model is trained so as to be able to determine, from spectrometric data derived from a sample that comprises a mixture of components (e.g. tissue types), an indication of a proportion in which the different components (e.g. tissue types) are (likely to be) present in the sample.
[0128] As shown in Figure 7A, at step 701, first reference spectrometric data representing a first component of the mixture, and second reference spectrometric data representing a second component of the mixture, is obtained. In embodiments, the first spectrometric data is derived from samples comprising (only) the first component, and the second spectrometric data is derived from samples comprising (only) the second component, e.g. substantially as described above with reference to Figure 4. At step 702, probability distributions are fitted to the first and second reference spectrometric data. In embodiments, Gaussian distributions are fitted to the data, e.g. substantially as described above with reference to Figure 5. Other distributions, such as the generalised Cauchy distribution described above, could be fitted to the data.
[0129] At step 703, training spectrometric data is generated using the fitted probability distributions. In embodiments, the fitted probability distributions are sampled (using a MCMC (Markov chain Monte Carlo) method) to generate “synthetic spectra” representing spectrometric data derived from a mixed sample. Step 703 may optionally comprise peak shape convolution (e.g. to provide appropriate resolution), and / or intensity scaling, and / or resampling (e.g. to restore instrument statistics, e.g. Poisson statistics for ToF). Then at step 304, the “synthetic” training spectrometric data is used as training data to train a machine learning model.
[0130] Although Figure 7A illustrates two components, one or more further components may be considered in a corresponding manner.
[0131] The approach according to various embodiments allows the generation of training data representative of mixed samples, which may otherwise be difficult to obtain.
[0132] As shown in Figure 7B, after training, the machine learning model can be used to characterise a sample. As shown in Figure 7B, at step 711, sample spectrometric data is obtained from a sample, and then at step 712, the sample spectrometric data is input to the trained machine learning model, which produces an output characterising the sample.
[0133] Figure 8 shows an artificial neural network trained according to the method described above, according to embodiments. The artificial neural network may be any suitable type of neural network, but in the present embodiment, the neural network is a convolutional neural network (CNN). The CNN comprises a number of layers 801-812 which operate one after the other, such that the output data from one layer is used as the input data for a next layer.
[0134] The CNN shown in Figure 8 comprises a first layer 801. The first layer 801 may receive an input spectrometric data array, apply a convolution and / or activation and / or pooling operation to the input spectrometric data array and pass a resulting output data array on to the next layer 802 of the neural network. The next layer 802 may then apply a convolution and / or activation and / or pooling operation and pass a resulting output data array on to the next layer 803, and so on. An activation operation can comprise any suitable activation function, such as a Sigmoid, Scaled Exponential Linear Unit (SELU) function and / or a bias. In the present embodiment a Rectified Linear Unit (ReLU) activation function is used.
[0135] The final layer 812 may produce a final output which may comprise a useful output. In the present embodiment, the output is a probability that the input spectrometric data array represents a sample comprising greater than a threshold percentage of breast tumour tissue. Other outputs would be possible.
[0136] Although Figure 8 shows a certain number of layers arranged in a certain topology, the neural network may comprise fewer or more layers if desired (and may also or instead comprise other layers which operate in a different manner to the layers described herein), and other topologies may be used, if desired.
[0137] Synthetic spectra were generated by, for each datapoint / of a training spectrum: sampling a Gaussian distribution having mean and standard deviation fi, On determined as shown in Figure 5A to determine a “synthetic” normal breast tissue intensity value; sampling a Gaussian distribution having mean and standard deviation g ogidetermined as shown in Figure 5B to determine a “synthetic” breast tumour tissue intensity value; and summing the intensity values according to a weighted sum, with the weight being selected to represent a particular proportion in which the breast tumour tissue is present.
[0138] Figure 9A illustrate results of a neural network trained using the synthetic spectra to output a probability that the input spectrometric data array represents a sample comprising greater than 5% breast tumour tissue. As shown in Figure 9A, the neural network was able to accurately identify samples comprising greater than 5% breast tumour tissue. Good results were also found for neural networks trained using the synthetic spectra to output a probability that the input spectrometric data array represents a sample comprising greater than 10% breast tumour tissue (Figure 9B) and 20% breast tumour tissue (Figure 90).
[0139] Although the embodiment described above involves training a machine learning mode using synthesised data, it would be possible to build and / or test other types of model using data synthesised in the same manner.
[0140] Although the embodiments described above generally involve analysis of mass spectral data, it will be appreciated that other spectral data, such as ion mobility spectral data, may be analysed. Although the embodiments described above generally involve analysis of a sample that may comprise a mixture of tissue types (normal and cancerous), it will be appreciated that other mixed samples, such as an adulterated food sample, may be analysed.
[0141] The foregoing detailed description has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the technology to the precise form disclosed. Many modifications and variations are possible in the light of the above teaching. The described embodiments were chosen in order to best explain the principles of the technology and its practical application, to thereby enable others skilled in the art to best utilise the technology in various embodiments and with various modifications as are suited to the particular use contemplated. It is intended that the scope be defined by the claims appended hereto.
Claims
166403 / 02Claims1. A method of spectrometric analysis comprising: providing an expression for the probability that spectrometric data represents a sample that comprises a mixture of at least a first component and a second component in a particular proportion; obtaining sample spectrometric data from a sample that may comprise a mixture of at least the first component and the second component; and using the expression and the sample spectrometric data to determine information relating to a proportion in which at least the first component and / or the second component is present in the sample.
2. The method of claim 1, wherein the expression is provided by: providing a prior probability distribution of a proportion; providing a likelihood of spectrometric data given the proportion; and determining a joint probability distribution of the spectrometric data and the proportion from the prior probability distribution and the likelihood.
3. The method of claim 1 or 2, wherein the expression is, or is based on equation (10) or equation (4):
4. The method of any preceding claim, comprising: providing first information indicative of statistical properties of spectrometric data representing samples comprising the first component; providing second information indicative of statistical properties of spectrometric data representing samples comprising the second component; and using the expression, the sample spectrometric data, the first information and the second information to determine the information relating to a proportion in which at least the first component and / or the second component is present in the sample.
5. The method of claim 4, wherein the first information is provided by: obtaining first reference spectrometric data from first reference samples comprising the first component; and fitting one or more first probability distributions to the first reference spectrometric data; and wherein the second information is provided by: obtaining second reference spectrometric data from second reference samples comprising the second component; and fitting one or more second probability distributions to the second reference spectrometric data.
6. A method of spectrometric analysis comprising: providing a machine learning model for determining, from spectrometric data representing a sample that comprises a mixture of at least a first component and a second component, information relating to a proportion in which at least the first component and / or the second component is present in the sample; obtaining sample spectrometric data from a sample that may comprise a mixture of at least the first component and the second component; and using the machine learning model and the sample spectrometric data to determine information relating to a proportion in which at least the first component and / or the second component is present in the sample.
7. The method of any preceding claim, wherein the information relating to a proportion indicates one or more of: (i) a value of the proportion; (ii) an uncertainty on a value of the proportion; (iii) a confidence interval on a value of the proportion; (iv) a probability of the proportion being less than, and / or equal to, and / or greater than, a threshold value; (v) a probability of the proportion being within a range of values; (vi) a probability of at least the first component and / or the second component being present in the sample in a proportion less than, and / or equal to, and / or greater than, a threshold value; (vii) a probability of at least the first component and / or the second component being present in the sample in a proportion within a range of values; and (viii) a class of plural classes representing different values of the proportion.
8. A method of providing a model for characterising a sample that may comprise a mixture of at least a first component and a second component using spectrometric data representing the sample, the method comprising: obtaining first reference spectrometric data from first reference samples comprising the first component; obtaining second reference spectrometric data from second reference samples comprising the second component; fitting one or more first probability distributions to the first reference spectrometric data; fitting one or more second probability distributions to the second reference spectrometric data; using the one or more first probability distributions and the one or more second probability distributions to synthesise data representative of spectrometric data derived from a sample comprising a mixture of at least the first component and the second component; and using the synthesised data to build and / or test a model for characterising a sample that may comprise a mixture of at least the first component and the second component using spectrometric data representing the sample.
9. The method of claim 8, wherein the model is a machine learning model, and the method comprises using the synthesised data to train the machine learning model to be able to characterise a sample that may comprise a mixture of at least the first component and the second component using spectrometric data representing the sample.
10. A method of spectrometric analysis comprising characterising a sample using a model provided according to claim 8 or 9.
11. The method of claim 5, 8, 9 or 10, wherein fitting one or more probability distributions comprises extracting plural features from spectrometric data, and fitting a respective probability distribution to data for each feature.
12. The method of claim 5, 8, 9, 10 or 11, wherein fitting one or more probability distributions comprises fitting one or more Gaussian distributions.
13. The method of any preceding claim, wherein the first component is a first tissue type, and the second component is a second tissue type.
14. The method of claim 13, wherein the first tissue type is normal tissue, and the second tissue type is tumour tissue.
15. The method of any one of claims 1 to 12, wherein the second component is a contaminant or adulterant.
16. The method of any preceding claim, wherein the spectrometric data comprises mass spectrometric data and / or ion mobility spectrometric data.
17. A method of operating an analytical instrument, comprising a method as claimed in any preceding claim, optionally wherein the analytical instrument is a mass spectrometer and / or ion mobility spectrometer.
18. A spectrometric analysis system comprising control circuitry arranged and adapted to perform a method as claimed in any one of claims 1 to 17.
19. An analytical instrument comprising control circuitry arranged and adapted to perform a method as claimed in any one of claims 1 to 17, optionally wherein the analytical instrument is a mass spectrometer and / or ion mobility spectrometer.
20. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out a method as claimed in any one of claims 1 to 17.