Control test for spectroscopic measurement
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-28
- Publication Date
- 2026-08-13
Smart Images

Figure GB2026050104_13082026_PF_FP_ABST
Abstract
Description
[0001] Control Test for Spectroscopic Measurement
[0002] FIELD OF INVENTION
[0003] The present disclosure is in the field of spectroscopy and relates particularly to methods and apparatus for control testing for spectroscopy, for quality testing of a measured spectrum, and particularly but not exclusively to control and quality testing of spectroscopic equipment, methods of control sample preparation, and related testing methods and protocols that involve or make use of controls and testing the quality of a measured spectrum. An example of the present disclosure relates to control testing and the quality testing of spectroscopic diagnostic equipment such as in determining diagnostic information in blood, such as cancer.
[0004] BACKGROUND TO INVENTION
[0005] In vitro diagnostic devices must be validated to ensure that the diagnostic test result is accurate and reliable. Clinical diagnostic controls are typically used with this process. By analysing samples with known expected results and comparing them to test results, control measurements are used to monitor the performance of equipment, reagents, and analysts. This ensures compliance with regulatory requirements as well as aiding in the detection and resolution of any hardware or software issues. Diagnostic controls and quality test processes also provide confidence in test results for the clinicians and patients benefitting from the in vitro device.
[0006] Diagnostic devices may measure, or quantify, specific target biomarkers. This targeted analysis, where the compound is known and can therefore be specifically identified and quantified, can be routinely checked using quality control samples to continuously monitor the performance of the analytical method. These control samples will often comprise a blank matrix sample spiked with known concentrations of the target analyte to determine whether the method is reliable and accurate within the expected concentration range. These can be considered ‘positive’ controls. Conversely, negative control samples can be provided to ensure that the system is able to provide a true negative result.
[0007] On the other hand, diagnostic devices may be non-specific, or ‘biomarker agnostic’, also known as non-targeted analysis. Non-targeted analysis, such as infrared (IR) and Raman spectroscopy and mass spectrometry, to name a few, investigates
[0008] 55763132-1multiple analytes, of which most are often unknown. Non-targeted analysis is common for the analysis of biofluids, such as blood, as they are complex, multi-component sample types. When such samples contain hundreds to thousands of components with varying concentrations and structural types, most of which are unknown prior to analysis, specific quality control samples like those used in targeted analysis are not appropriate for the quality control (QC) of non-targeted analysis. Even if all these components were known beforehand, preparing multiple QC samples across the concentration range for numerous components would not be realistic.
[0009] Sangsteret al (T. Sangster, H. Major, R. Plumb, A. J. Wilson, and I. D. Wilson, A pragmatic and readily implemented quality control strategy for HPLC-MS and GC-MS-based metabonomic analysis, Analyst, 2006, 131, 1075-1078) proposed a QC method for non-targeted analysis. They used pooled quality control samples, which consisted of a combination of aliquots from every sample due for analysis to create a pooled sample. This pooled sample is then analysed regularly throughout sample analysis. Using Principal Component Analysis (PCA), the QC data would be expected to cluster together to demonstrate there were no time-related trends within sample analysis. Examples of applications of this method can be found in the context of mass spectrometry metabolomics as well as IR spectroscopic applications within the food industry and healthcare. For example:
[0010] E. Achten, D. Schiitz, M. Fischer, C. Fauhl-Hassek, J. Riedl, and B. Horn Classification of grain maize (Zea mays L.) from different geographical origins with FTIR spectroscopy-a suitable Analytical tool for feed authentication?, Food Analytical Methods, 2019, 12, 2172-2184.
[0011] B. Horn, S. Esslinger, M. Pfister, C. Fauhl-Hassek, and J. Riedl, Non-targeted detection of paprika adulteration using mid-infrared spectroscopy and one-class classification - is it data preprocessing that makes the performance?, Food Chemistry, 2018, 257, 112-119.
[0012] T. Nietner, M. Pfister, M. A. Glomb, and C. Fauhl-Hassek, Authentication of the botanical and geographical origin of distillers dried grains and solubles (DDGS) by FTIR spectroscopy, Journal of Agricultural and Food Chemistry, 2013, 61 , 7225-7233.
[0013] M. Huber, K. W. Kepesidis, L. Voronina, F. Fleischmann, E. Fill, J. Hermann, I. Koch, K. Milger-Kneidinger, T. Kolben, G. B. Schulz, F. Jokisch, J. Behr, N. Harbeck, M. Reiser, C. Stief, F. Krausz, and M. Zigman, Infrared molecular fingerprinting of bloodbased liquid biopsies for the detection of cancer, eLife, 2021, 10, e68758.
[0014] 55763132-1M. Huber, K. V. Kepesidis, L. Voronina, M. Bozic, M. Trubetskov, N. Harbeck, F. Krausz, and M. Zigman, Stability of person-specific blood-based infrared molecular fingerprints opens up prospects for health monitoring, Nature Communications, 2021, 12, 1511.
[0015] The manual process of analysing and comparing QC data with sample data using PCA is time-consuming. Also, it requires the user to determine how closely the data must be clustered in order to be deemed acceptable, which can vary among analysts as it is subjective. Additionally, the use of a pooled QC sample that is created during the analysis from aliquots of the samples being analysed is not feasible in some settings, e.g. clinical applications where patient samples could be entering the clinic at multiple time points throughout the day, rather than in strict batches.
[0016] An example of non-targeted analysis is provided by the Dxcover (RTM) Liquid Biopsy Platform (by Dxcover Limited, UK), a novel blood serum test for the early detection of brain cancer. The system is an Attenuated Total Reflectance-Fourier Transform Infrared (ATR-FTIR) spectroscopy platform, coupled with a diagnostic machine learning model, that generates a disease prediction that can be used to inform clinical decisions. The qualitative, non-targeted approach tests for the signals indicative of cancer from blood serum samples obtained from patients. It is deemed non-targeted as it analyses the entirety of the human serum, which in itself possesses a range of biological markers which contribute to a ‘spectral signal’ indicative of the presence of cancer.
[0017] The presence of diagnostic information in blood is not a recent discovery. Clinically established blood-based tests for cancer, such as carcinoembryonic antigen (CEA) or prostate specific antigen (PSA), have been utilised in medicine for many decades. For example: A. Lokshin, R. C. Bast, and K. Rodland, Circulating Cancer Biomarkers, Cancers, 2021, 13, 802.
[0018] The emergence of liquid biopsy technologies has shed light on the diagnostic utility of novel biomarkers in blood samples. These recent developments have extended beyond simple protein detection, to highly sophisticated genetic testing that can isolate incredibly low levels of diagnostic markers, such as tumour-derived DNA or circulating tumour cells. For example: M. M. J. Bauman, S. M. Bouchal, D. D. Monie, A. Aibaidula, R. Singh, and I. F. Parney, Strategies, considerations, and recent advancements in the development of liquid biopsy for glioblastoma: a step towards individualized medicine in glioblastoma, Neurosurgical Focus, 2022, 53, E14.
[0019] 55763132-1The Dxcover (RTM) Liquid Biopsy test measures the infrared absorption of all molecules contained within blood serum, including the following molecular classes: proteins (e.g., albumin, globulin), lipids (e.g., cholesterol), carbohydrates (e.g., glucose), and phosphate. The Dxcover (RTM) Liquid Biopsy test does not measure one specific biomarker in the traditional sense of an analyte, but rather the molecular profile of the blood serum in the form of a spectroscopic signature. Whilst areas of the spectrum may be tentatively associated with known biomolecules, it is not a simple binary classification of the presence of individual biomarkers, but rather the overall signature of the patient profile. The clinically relevant information in the spectral signature likely originates from two sources: (i) the tumour itself, and (ii) the patient’s immune response to the tumour. The spectroscopic signature can only be interpreted by computational methods due to the complexity of the signal coded into the infrared spectrum. Machine learning algorithms can be used to learn patterns within the dataset that are not discernible by eye.
[0020] The Dxcover (RTM) Liquid Biopsy Platform is based upon both hardware and software components. The hardware items include an ATR-FTIR spectroscopy system, allowing high-throughput, automated, and efficient processing of serum samples in a clinical setting. The software components analyse the spectral output generated by the hardware and generates a disease prediction using a diagnostic model based on machine learning analysis.
[0021] The present inventors have appreciated some disadvantages of known control, and quality testing methods and hardware, and have recognised possible improvements thereto.
[0022] SUMMARY OF INVENTION
[0023] According to a first aspect of the disclosure, there is provided a method of control testing for a non-targeted analysis spectroscopic measurement, the method comprising:
[0024] (i) carrying out a first mode of control testing a spectrometer, independent of a subsequent non-targeted analysis diagnostic measurement to be carried out on a test sample by the spectrometer, the first mode comprising:
[0025] providing a first, reference sample to the spectrometer, wherein the first, reference sample is configured to be independent of the test sample;
[0026] obtaining, by a processor element, a first measured spectrum of the first, reference sample;
[0027] determining if the first measured spectrum is of acceptable quality; and / or
[0028] 55763132-1(ii) carrying out a second mode of negative control testing in respect of the subsequent diagnostic measurement to be carried out on the test sample by the spectrometer, the second mode comprising:
[0029] providing a second, negative control sample to the spectrometer, wherein the second sample is configured to produce a different measured spectrum to a measured spectrum of the test sample;
[0030] obtaining, by the processor element, a second measured spectrum of the second, negative control sample;
[0031] determining if the second measured spectrum is of acceptable quality; and / or (iii) carrying out a third mode of positive control testing in respect of the subsequent diagnostic measurement, the third mode comprising:
[0032] providing a third, positive control sample to the spectrometer, wherein the third sample is configured to produce a same or similar measured spectrum to the measured spectrum of the test sample;
[0033] obtaining, by the processor element, a third measured spectrum of the third, positive control sample;
[0034] determining if the third measured spectrum is of acceptable quality.
[0035] To avoid any confusion with the traditional meaning of positive and negative controls, the term “negative control” in this context means a control sample that will produce a spectrum that is deemed to be sufficiently different to an expected measured spectrum of the test sample. The term “positive control” in this context means a control sample that will produce a spectrum that is sufficiently close to or the same as an expected measured spectrum of the test sample. That is, the use of positive and negative controls does not imply any diagnostic result.
[0036] The method may comprise carrying out the subsequent non targeted analysis diagnostic measurement of the test sample. The non-targeted analysis may be a nonspecific spectroscopic measurement in which one or more or a plurality of unknown analytes are to be measured.
[0037] The processor element may be configured to determine the quality of the first measured spectrum, the second measured spectrum and / or the third measured spectrum. The / each step of determining the quality of the first measured spectrum, the second measured spectrum, and / or the third measured spectrum may comprise:
[0038] identifying, by the processor element, at least one spectrum parameter of the measured spectrum; and
[0039] 55763132-1comparing, by the processor element, the at least one spectrum parameter against at least one reference parameter.
[0040] Determining the quality of the, or each measured spectrum by the processor element may be based at least in part on the comparison of the at least one spectrum parameter to the at least one reference parameter.
[0041] The method may comprise generating, by the processor element, a quality score of the first measured spectrum, and / or the second measured spectrum, and / or the third measured spectrum. The quality score may be one or more discrete values, such as a binary pass or fail, 1 or 0, or the like, or a scored range of values, such as a range of percentage values or a grading system such as poor, good, acceptable, excellent, or the like. It will be understood that other quality scores or indicators may be used.
[0042] The quality of the measured spectrum may mean the quality of at least a part of the measured spectrum, or substantially the whole measured spectrum.
[0043] The processor element may be or may be part of a computing device, the spectrometer, a spectroscopy system, a measurement apparatus, a measurement apparatus comprising the spectrometer, a distributed computing system such as a cloud computing system, or the like. The processor element may be implemented virtually, such as using cloud computing or an emulator. The processor element may be integrated within or connected to the spectrometer.
[0044] The, or each measured spectrum of the first aspect may be obtained by the same spectrometer.
[0045] The spectrometer may be an infrared (IR) spectrometer, a Fourier Transform IR (FTIR)-based spectrometer, an Attenuated Total Reflectance-Fourier Transform Infrared (ATR-FTIR) spectrometer, a vibration-based spectrometer such as Raman spectroscopy, or any suitable spectrometer for obtaining a measured spectrum from a sample. The non-targeted analysis of the test sample may be an IR or FTIR or ATR-FTIR non-targeted analysis, which may include any suitable processing or analysis steps. The spectrometer may be configured for any spectroscopic technique and the IR and vibrational examples are provided as examples and are by no means essential to the invention.
[0046] The spectrometer is configured for carrying out diagnostic measurements. The spectrometer is configured for carrying out non-targeted analysis diagnostic measurements. The spectrometer may be configured for carrying out diagnostic measurements, and the method of control testing may be a control test to determine if such diagnostic measurements will be, or are likely to be valid. The spectrometer may
[0047] 55763132-1be configured to carry out diagnostic measurements using one or more artificial intelligence or machine learning algorithms or processes to analyse the measured spectrum of the test sample to measure one or more analytes and in some examples to diagnose or infer one or more health conditions, such as cancer, or the like. The spectrometer may be configured for performing diagnostic measurements on human or animal biological material, which may be blood, plasma or serum, or other biological material.
[0048] The method of control testing may be a quality test to determine if the measurement, diagnosis and / or or inference in respect of the test sample, when performed by the spectrometer, will be valid because the measured spectrum will be of sufficient quality. The method of control testing may include a numerical and / or predictive quality test of the first measured spectrum and / or the second measured spectrum and / or the third measured spectrum to determine if a subsequent diagnostic measurement will be or is likely to be valid.
[0049] The spectrometer may be configured to measure an intensity (e.g. absorbance or transmission) versus frequency (e.g. in wavenumbers), of the sample to be measured thereby. In some examples, the spectrometer may be configured to obtain a measured spectrum including frequencies corresponding with at least some or all molecules contained within human / animal blood, plasma, or serum. The spectrometer may be configured to obtain a measured spectrum that includes a signal indicative of at least one of: proteins (e.g., albumin, globulin), lipids (e.g., cholesterol), carbohydrates (e.g., glucose), deoxyribonucleic acid (DNA), ribonucleic acid (RNA), sugars, drugs, metabolites, phosphate, and / or any other component, in the human / animal blood, plasma or serum.
[0050] The method may comprise providing a sample to the spectrometer for each repetition, or each mode, of the method.
[0051] The method may comprise one or more sample preparation steps for the, or each sample.
[0052] In use, the processor element may be remote from the sample. The processor element may be integrated with, connected to, or remote from the spectrometer. The spectrometer may comprise the processor element. When remote from the spectrometer, the method may comprise obtaining the measured spectrum or spectra from the spectrometer to the processor element.
[0053] In some examples, the step of determining the quality of the first measured spectrum, the second measured spectrum, and / or the third measured spectrum may be
[0054] 55763132-1carried out automatically by the processor element. The spectrometer and the processor element may be configured to run the or each quality determination step, which may be automatically. Running the quality determination step or steps may be in response to user input provided to the processor element or in some examples to at least one of the processor element and the spectrometer. The processor element may be programmed to run the quality determination step(s), or in some examples, the spectrometer and the processor element may be programmed or otherwise configured to run the quality determination step(s).
[0055] The processor element and the spectrometer may form a system configured to carry out control testing, and in some examples, a system configured to determine the quality of measured control spectra and to carry out a method of diagnosis on the test sample. In some examples, the spectrometer and processor element can run a control test on one or more or a plurality of control samples, and then a diagnostic method on the test sample, so that the validity of the diagnostic method can be predicted before it is run.
[0056] The processor element may comprise a user interface. The spectrometer may comprise a user interface. The spectrometer user interface may be the user interface of the processor element.
[0057] The processor element may comprise a display element. The spectrometer may comprise a display element. The display element of the spectrometer may be the display of the processor element.
[0058] The method of control testing may comprise a first user input step to start the, or each mode of the method. The first user input step may begin the step of obtaining a measured spectrum of the, or each mode. For the, or each mode, the method may comprise determining the quality of the measured spectrum by the processor element without any further user input. The method of control testing may be an automated method carried out by the processor element from the first user input to the determination of the quality of the measured spectrum for the, or each mode. In some examples, the method of control testing may be at least partially automated by the processor element and the spectrometer.
[0059] The method may comprise carrying out at least some of the steps of the method again to obtain a further spectrum or further spectra from a control sample and determining the quality of each measured spectrum thereby obtained. The method may be repeatable each time by a further user input step.
[0060] 55763132-1The processor element and / or the spectrometer may be configured to prevent a diagnostic measurement of the test sample unless the first measured spectrum is of acceptable quality, and / or unless the second measured spectrum is of acceptable quality, and / or unless the third measured spectrum is of acceptable quality. The processor element and / or the spectrometer may be configured to prompt a repeat of the first mode and / or the second mode and / or the third mode if at least one mode returns a spectrum of unacceptable quality, before permitting a diagnostic measurement of the test sample to be carried out.
[0061] The / each first user input step may be an activation provided to a user input device, which may take any suitable form, such as button(s), touch screen interface, keyboard, mouse, or the like.
[0062] The method of control testing may comprise carrying out a plurality of modes of operation. The method of quality testing may comprise carrying out at least two of the first mode, the second mode and the third mode. The method may comprise carrying out one or more, or a plurality of runs of the first mode, one or more or a plurality of runs of the second mode, and / or one or more or a plurality of runs of the third mode. The method may comprise running the first mode, the second mode and / or the third mode on a single substrate comprising one or more samples, or a plurality of samples on a plurality of substrates, or one substrate and sample for each mode. The method may comprise running the first, second and third modes.
[0063] The first mode may be a method of quality testing hardware, such as the spectrometer, using the first, reference sample. The second mode may be a method of quality testing using the second, negative control sample. The third mode may be a method of quality testing using the third, positive control sample. The method may comprise other modes than those listed here.
[0064] The first mode may be a hardware and / or software verification mode for determining if the spectrometer is working as expected. The first mode may be for determining that one or more further components of a diagnostic system are working as expected in addition to the control testing of the spectrometer. The first mode may be for determining that the spectrometer is working as expected independent of a biological sample. The first mode may be for determining that the spectrometer is working as expected, which may be a test that is independent of biological material, such as reagents, or the like. The first mode is independent of a subsequent non-targeted analysis diagnostic measurement to be carried out on the test sample by the spectrometer and the first mode may be independent of the second mode and / or the
[0065] 55763132-1third mode. In this example, the performance of the first mode, including the selection of the properties of the first, reference sample does not depend on the second and / or third modes. In some examples, the first mode and the first reference sample are independent of the second mode, the third mode, and the subsequent measurement to be carried out on the test sample.
[0066] The spectrometer and / or the processor element may be configured to not permit the second mode and / or the third mode to be carried out unless the first mode returns an acceptable first measured spectrum.
[0067] The second mode may be a control test for determining if a negative control sample will be identified as such by the spectrometer because the measured spectrum is not close enough to an expected measured spectrum of the test sample. The spectrometer and / or the processor element may be configured to provide a fail result for a successful run of the second mode because the second measured spectrum is sufficiently different from an expected spectrum of the test sample. The spectrometer and / or the processor element may be configured to provide a pass result if the second measured spectrum is sufficiently similar to an expected spectrum of the test sample. The second mode may be for determining if a user can prepare a true negative sample that will not be identified by the spectrometer as being comparable to an expected measured spectrum of the test sample, and a successful result may be a fail being obtained which indicates correct preparation of a negative control.
[0068] The third mode may be a quality test for determining if a positive control sample will be identified as such by the spectrometer because the measured spectrum is close enough to an expected measured spectrum of the test sample. The spectrometer and / or the processor element may be configured to provide a pass result if the third measured spectrum is sufficiently similar to an expected measured spectrum of the test sample. The third mode may be for determining if a user can prepare a true positive sample and is operating the spectrometer correctly to enable a measured spectrum to provide an expected positive match result for a positive control sample, such that it can be expected that the measurement of the test sample will be performed correctly.
[0069] The method may comprise running the first mode before the second mode and / or the third mode. The method may comprise running the second mode before the third mode or the third mode before the second mode. The method may comprise running the first mode then the second mode then the third mode. The method may comprise running the diagnostic method after at least one of the first, second and third modes have been run, or when at least two of, or all three modes have been run.
[0070] 55763132-1The method may comprise obtaining more than one spectrum for the, or each sample.
[0071] The measured spectrum may comprise a plurality of data points. The plurality of data points may each include an intensity value (e.g. absorbance or transmittance) for a frequency value (e.g. wavenumber). The plurality of data points may be distributed between a range of frequencies, which may be between about 0 cm-1and about 5,000 cm-1, or between about 500 cm-1and about 4,000 cm-1, or between about 1 ,000 cm-1and about 3,500 cm-1.
[0072] The first sample, the second sample, and the third sample are control samples. The test sample may be a diagnostic test sample, which may be provided to the spectrometer after the method of control testing. The test sample may comprise biological material, such as human / animal blood, plasma or serum, or other biological material.
[0073] The method of control testing may be carried out using a plurality of, or three or more samples for the, or each mode and each sample may be a control sample.
[0074] The first, reference sample may comprise no subject test samples, such as diagnostic test samples. A subject test sample means a sample from a human or animal for whom a test result is to be obtained from the test sample. The second, negative control sample may comprise no subject test samples or portion of subject test samples. The third, positive control sample may comprise no subject test samples or portion of subject test samples. The first, reference sample, the second negative control sample and the third, positive control sample may be devoid of subject test material intended for use with the test sample.
[0075] The sample may be located on / in a substrate, which may take any suitable form. The substrate may be a slide, chip, die, planar member, or the like, which may be silicon-based or made of silicon or comprise silicon. The substrate may comprise a housing, cover or the like for the slide, chip, die, or planar member. The substrate may be configured for use with the spectrometer, and in some examples may be a slide or substrate configured for use with an I R, FTIR, and / or ATR-FTIR spectrometer, or Raman Spectrometer, or the like. At least one of the first, second and third samples may be implemented on a common substrate.
[0076] The spectrometer may be configured to generate multiple spectra from a plurality of the samples. For example, by measuring a sample more than once, or sequentially measuring two or more samples, or in any suitable manner.
[0077] 55763132-1The first sample may be a non-biological sample. The first sample may be a blank or neutral control sample. The first sample may be a non-biological control sample. The first sample may be substantially devoid of any biological material. The first sample may be substantially devoid of any biological material present in the second sample and / or the third sample. The first sample may be substantially different to the test sample. The first sample may be substantially devoid of any biological material present in the test sample. The first sample may be configured to be substantially biologically inert. The first sample may be independent of the second sample and / or the third sample. The method may comprise selecting the first sample independent of the second and / or third sample material properties such that the first sample is independent therefrom. The first sample is independent of the test sample, and different therefrom. The first sample may comprise or may be a plastics or polymer-based sample, or may comprise or may be a plastic or polymer sample. The first sample may be a solid and / or liquid sample. The first sample may comprise or may be a resin sample. The first sample may be a synthetic sample. The first sample may comprise or may be an epoxy-based sample, or may comprise or may be an epoxy sample, or may comprise an epoxy substance on a substrate. In some examples the first sample may be an epoxy resin material. The first sample may be located on or in any suitable substrate.
[0078] The first, reference sample may be configured to be re-usable. The first sample may be reusable over at least 24 hours, or at least 1 week, or at least 1 month, or at least 1 year or at least 2 years, or 2 years or less. The first, reference sample may be configured to be a stable reference sample.
[0079] The first, reference sample may be configured to produce one or more spectral peaks and / or troughs, or two or more spectral peaks and / or troughs, or four or more spectral peaks and / or troughs, or the like, when measured by the spectrometer. The first reference sample may be configured to produce at least one signal-free region, or two or more, or four or more.
[0080] The second and / or third samples may be biological samples.
[0081] The first, second and third samples may be different to each other.
[0082] The second sample may comprise human or animal biological material. The second sample may be different to the test sample. The second sample may comprise one or more components of and / or one or more comparable components to the third sample and / or to the test sample. The one or more components may be a protein. The one or more comparable components may include an animal component comparable to a human component. The second sample may comprise at least a portion, or a
[0083] 55763132-1comparable portion, of the third sample and / or the test sample. The component or comparable component may include a serum component, a plasma component, and / or a blood component. The component or comparable component may include human and / or animal component(s). The second sample may comprise or may be blood, plasma, serum, bovine serum albumin (BSA), a BSA solution, bovine plasma albumin, a bovine plasma albumin solution, or any suitable biological material. The second sample may be a liquid and / or a solid sample.
[0084] The second, negative control sample may be configured to be stable, optionally for at least 1 day, or at least 1 week, or at least 1 month, or for 1 month or less, or about 1 month, or between 1 day and 1 month, or between 1 week and 1 month, or any suitable period of time.
[0085] The method of control testing may comprise selecting or preparing the second, negative control sample based, at least in part, on one or more properties of the test sample.
[0086] The third sample may comprise human or animal biological material. The third sample may be different to the test sample. The third sample may comprise one or more components, or a plurality of components, or substantially all components of the test sample. The shared component(s) of the third sample and the test sample may be biological components, which may be components of blood, serum, and / or plasma. The test sample may comprise a carrier component, which may be blood, plasma or serum, and the third sample may comprise substantially the same carrier component as the test sample. The third sample and the test sample may be human samples, which in some examples may be human blood, plasma or serum samples. The third sample may be configured to match the carrier type of the test sample. The third sample may have at least one more component in common, or comparable component in common, with the test sample than the second sample, or 5 or more, 10 or more, 20 or more, or any suitable number of common components or comparable components. The third sample may comprise or may be blood, plasma, serum, lyophilised human pooled serum (LHPS), or a LHPS solution, lyophilised human pooled plasma, or a lyophilised human pooled plasma solution, or any suitable biological material. The third sample may be a liquid and / or a solid sample. The test sample may be a liquid and / or a solid sample. The test sample may have substantially the same volume as the third sample and / or the second sample.
[0087] The third, negative control sample may be configured to be stable, optionally for at least 1 day, or at least 1 week, or at least 1 month, or for 1 month or less, or about 1
[0088] 55763132-1month, or between 1 day and 1 month, or between 1 week and 1 month, or any suitable period of time.
[0089] The method of control testing may comprise selecting or preparing the third, positive control sample based, at least in part, on one or more properties of the test sample.
[0090] The, or at least some of the control samples may be devoid of reagents or added reagents. The first, reference sample may be devoid of reagents. The second, negative control sample may be devoid of reagents.
[0091] The at least one reference parameter may include one or more spectral parameters. The at least one reference parameter may include one or more parameters of a reference spectrum, and / or one or more parameters derived from, at least in part, a plurality of reference spectra.
[0092] The step of comparing the at least one measured spectrum parameter against the at least one reference parameter may include comparing two or more measured spectrum parameters against two or more reference parameters, optionally four or more, optionally 8 or more, optionally 15 or more, optionally 20 or more, optionally 40 or more, or at least 2, or at least 4, or at least 8, or at least 15, or at least 20, or at least 40.
[0093] The, or each reference parameter may be a numerical value or a numerical range. The numerical range may be open at one end, or be a closed range.
[0094] The at least one reference parameter may include one or more of the following: the intensity values for a range of frequencies;
[0095] an intensity value for a frequency;
[0096] a location of one or more expected peak values or trough values;
[0097] a peak intensity value or trough intensity value (maximum value or minimum value) at one or more frequencies;
[0098] a prominence of one or more peak intensity values or trough intensity values for a range of frequencies;
[0099] a signal to noise ratio of an intensity value, or intensity values of the spectrum, which may include the signal to noise ratio of a peak intensity value or trough intensity value against another part of the spectrum, such as a non-signal part or non-peak / trough part of the spectrum; and / or
[0100] a relative peak to trough value within a range of frequencies.
[0101] In some examples, the reference parameters may include multiple parameters of the same type. For example, intensity values for different ranges of frequencies.
[0102] 55763132-1A prominence may be a vertical distance between the peak intensity value and its lowest contour line, or between a trough intensity value and its highest contour line. It will be understood that the prominence can be defined in many different ways.
[0103] The, or each identified spectrum parameter of the measured spectrum may be selected by the processor to correspond substantially to at least one reference parameter.
[0104] The identified at least one spectrum parameter may be identified as contributing to the quality of the measured spectrum. This may be carried out by the processor element and / or by input to the processor element.
[0105] The at least one identified spectrum parameter may be identified based on the mode of operation and / or the sample that is to be, or has been, used to obtain the measured spectrum. The method may comprise pre-configuring or selecting information about the mode of operation, the test to be run, and / or the sample.
[0106] At least one, or some, or all of the reference parameters may be different between at least two of the first mode, the second mode and the third mode.
[0107] In at least one mode of control testing, the step of determining the quality of the measured spectrum may include using a numerical comparison between the at least one spectrum parameter and the at least one reference parameter.
[0108] The first mode may comprise numerically comparing the at least one identified spectrum parameter to the at least one reference parameter. The comparison may generate a binary quality measure (e.g. true or false), or a probabilistic measure, or any suitable quality measurement.
[0109] The numerical comparison may comprise at least one of: comparing a measured parameter to a reference parameter using an arithmetic operation, such as subtraction, or a vector operation such as obtaining an Euclidean norm of the difference of a measured spectrum parameter to a reference parameter, or the like.
[0110] At least one of the steps of: comparing the at least one spectrum parameter against the at least one reference parameter; and determining the quality of the measured spectrum may comprise using a machine learning algorithm. In at least one mode of control testing, the step of determining the quality of the measured spectrum may comprise using a machine learning algorithm. In at least one mode of control testing, the step of determining the quality of the measured spectrum may comprise using a numerical comparison method and in at least one mode of control testing, the step of determining the quality of the measured spectrum may comprise using a machine learning algorithm. The machine learning algorithm may be a trained model. It will be
[0111] 55763132-1understood that in at least some examples, the machine learning algorithm or trained model may be reconfigured, updated, modified, which may happen any suitable number of times or in any suitable way, such that in at least some embodiments the model or algorithm is not fixed but may be modified or optimised over time.
[0112] The processor element may comprise one or more quality testing machine learning algorithms for determining if a measured spectrum from a sample is of acceptable quality. The processor element may be operable to run or otherwise use the quality testing machine learning algorithm, which may include running the machine learning algorithm on a further device, such as a further computing device. The machine learning algorithm may be run in various ways, such as on a computing device that the processor element is part of, on another computing device, on an emulator or cloud server or the like, and in some examples different functions of the machine learning algorithm may be run on different devices. It will also be understood that the term machine learning algorithm does not exclude artificial intelligence, deep learning, or other techniques applicable to the present invention. It is also possible that more than one algorithm may be used. It is also envisaged that in some examples multiple runs of the machine learning algorithm may be carried out to obtain an average result or to allow for erroneous runs to be discarded.
[0113] The, some, or all of the steps ofdetermining the quality of the measured spectrum may comprise using the quality testing machine learning algorithm. The second mode and / or the third mode may comprise using the quality testing machine learning algorithm.
[0114] The quality testing machine learning algorithm may be a decision tree method, which may be an ensemble of decision trees. The machine learning algorithm may be a supervised learning algorithm. The decision tree method may be a random forest model, or any suitable algorithm. It will be understood from a reading of the specification that the quality testing machine learning algorithm may be any suitable machine learning algorithm that is trained to determine the quality of the measured spectrum, and it is not essential to use any particular form of machine learning. The specific examples of machine learning provided herein are purely for example to understand how the invention may be implemented.
[0115] The quality testing machine learning algorithm may be different to the diagnostic machine learning or artificial intelligence algorithm that may be used for diagnostic analysis or measurement of the test sample.
[0116] The processor element may be configured to use the quality testing machine learning algorithm for the second mode. The processor element may be configured to
[0117] 55763132-1use the machine learning algorithm for the third mode. The quality testing machine learning algorithm may be trained for one or more modes of operation, or two or more modes of operation, or two modes, of the method of control testing. The training may depend on the intended control sample, or control samples, to be used.
[0118] The at least one reference parameter may be reference machine learning parameter(s) when machine learning is used.
[0119] The first mode may use at least one reference parameter, and the second mode and / or the third mode may use machine learning and at least one reference machine learning parameter. The first mode may be a numerical comparison test and the second mode may be a machine learning prediction or classification test.
[0120] The, or each reference machine learning parameter may include one or more spectral parameters. The at least one reference machine learning parameter may include one or more parameters of a reference spectrum, and / or one or more parameters derived from, at least in part, a plurality of reference spectra.
[0121] The, or each reference machine learning parameter may be a numerical value or a numerical range. The numerical range may be open at one end, or be a closed range.
[0122] The step of comparing the at least one spectrum parameter against the at least one reference machine learning parameter may include comparing two or more spectrum parameters against two or more reference parameters, optionally four or more, optionally 8 or more, optionally 15 or more, optionally 30 or more, optionally 50 or more, optionally 100 or more, or at least 2, or at least 4, or at least 8, or at least 15, or at least 30, or at least 50, or at least 100.
[0123] The quality testing machine learning algorithm may use the comparison of measured parameters to reference parameters to determine the quality of the measured spectrum or spectra using any suitable machine learning technique. It will be understood that the comparison of the measured parameters to the reference machine learning parameters may be carried out in many ways, and it may not involve a direct comparison, but the machine learning algorithm may use the reference machine learning parameters to learn how to determine the quality of a measured spectrum.
[0124] It will be understood that the machine learning algorithm may be configured to determine the quality of a measured spectrum without directly looking up a reference parameter, but in this example, the comparing of the at least one spectrum parameter against at least one reference machine learning parameter still effectively takes place as the machine learning algorithm will need to make decisions based on reference parameters.
[0125] 55763132-1In some examples, the machine learning algorithm may be trained on spectra, rather than the reference parameters.
[0126] The at least one reference machine learning parameter may include one or more of:
[0127] a location of a peak intensity value or trough intensity value within a range of frequencies;
[0128] a location of a maximum or minimum intensity value within a range of frequencies;
[0129] a value of a peak or trough for a range of frequencies;
[0130] a location of a peak intensity value or trough intensity value within a region of the spectrum associated with a substance (e.g. within an absorbance band of the substance);
[0131] a location of a local or global peak or global trough within a range of frequencies; the intensity value of the local or global peak / trough;
[0132] spectral intensity values for a range of frequencies;
[0133] a transformation parameter used to apply a transformation or correction to the measured spectrum or spectra, or to the measured spectrum parameter(s);
[0134] a ratio of the intensity value of a local maximum within a range of frequencies to a local minimum within a range of frequencies, wherein the ranges may be different or substantially the same;
[0135] a signal to noise ratio, which may include the signal to noise ratio of a peak intensity value or trough intensity value against another part of the spectrum, such as a non-signal part or non-peak / trough part of the spectrum;
[0136] a prominence of one or more peak values or trough values;
[0137] a sum of prominences of a local minima and maxima within a range of frequencies;
[0138] a width of a peak or trough for a range of frequencies, which may be for a substance band.
[0139] The reference machine learning parameter(s) may include any features or options of the reference parameters and vice versa.
[0140] Any suitable range of frequencies may be selected to correspond with an absorbance band of a substance.
[0141] For the signal to noise ratio, the signal portion may include an absorbance band of a substance or a range of values for a range of frequencies, and the noise portion may
[0142] 55763132-1be outside of the absorbance band, or at a non-peak / non-trough section, or at a nonsignal section.
[0143] The transformation parameter may be one or more parameters of a function for correcting a slope of the measured spectrum. In this example, the transformation applied to the measured spectrum may be compared with the reference value or range of values that are deemed acceptable for this parameter to obtain a machine learning reference parameter.
[0144] The substance region or band may be any suitable substance band used with spectroscopic analysis, such as an amide band, such as amide I, or amide II. These are examples and are not intended as being essential to the invention.
[0145] When using the quality testing machine learning algorithm, in some examples the measured spectrum may be quality tested through the algorithm having been trained on reference spectra, rather than reference parameters, but in other examples the machine learning algorithm may use a comparison of the at least one spectrum parameter to the at least one reference machine learning parameter. When using the quality testing machine learning algorithm, the step of comparing the at least one spectrum parameter to the at least one reference machine learning parameter may comprise measuring one or more of:
[0146] a peak or trough x-shift (frequency shift) of the measured spectrum parameter against the reference machine learning parameter;
[0147] a peak or trough y-shift (difference in intensity) of the measured spectrum parameter against the reference machine learning parameter;
[0148] a numerical comparison or arithmetic operation between the measured spectrum parameter and the reference machine learning parameter;
[0149] a calculation of the degree of similarity to expected value(s), which may be a percentage value or the like;
[0150] comparison to mean values, which may be a mean reference parameter obtained from a plurality of reference parameters, where the mean may be for a range of frequencies;
[0151] an area difference to the mean, which may be for a range of frequencies; a cosine distance to the mean, which may be for a range of frequencies;
[0152] a 1-norm distance to the mean, which may be for a range of frequencies; the number of intensity values that are outside of n standard deviations from the mean, where n can be e.g. 1, 2, 3, or any suitable value;
[0153] 55763132-1a ratiometric difference between a local maximum and a local minimum with a range of frequencies;
[0154] a sum of two or more prominences of local minimum / minima and local maximum / maxima in a range of frequencies;
[0155] determination if a parameter is above or below a threshold value, such as assigning a “1” if the width of an intensity peak is below a threshold width, and a “0” if above the threshold width.
[0156] In these examples, the mean may be a mean value of a plurality of reference spectra, which effectively creates a mean reference spectrum for identifying the similarity of a measured spectrum to a mean reference spectrum. As detailed above, the mean may be for a range of frequencies, such as a part of the mean reference spectrum.
[0157] The, or each identified spectrum parameter of the measured spectrum may be selected by the processor element to correspond substantially to at least one reference machine learning parameter.
[0158] At least one, or some, or all of the reference machine learning parameters may be different between the second mode and the third mode.
[0159] The method may comprise carrying out one or more runs of the machine learning comparison between the at least one spectrum parameter and the at least one reference machine learning parameter.
[0160] The second mode and / or the third mode may comprise comparing the at least one identified spectrum parameter to the at least one reference machine learning parameter to determine the quality score. The quality score may be one or more discrete values, such as a binary pass or fail, 1 or 0, or the like, or a scored range of values, such as a range of percentage values or a grading system such as poor, good, acceptable, excellent, or the like. It will be understood that other quality scores or indicators may be used.
[0161] The step of identifying the at least one spectrum parameter of the measured spectrum may comprise selecting values from the measured spectrum that correspond to the at least one reference parameter and / or the at least one machine learning reference parameter.
[0162] The identified spectrum parameters of the measured spectrum may each correspond to a reference parameter and / or a reference machine learning parameter.
[0163] The method may include providing the at least one spectrum parameter to the machine learning algorithm.
[0164] 55763132-1The method may comprise training the machine learning algorithm. The machine learning algorithm may be a trained algorithm. The machine learning algorithm may be trained on one or more, or a plurality of reference spectrum / spectra, which may form a training data set.
[0165] The machine learning algorithm may be trained for determining the quality of different expected results, which may include the quality of an expected negative result (is not close to an expected test sample spectrum) or an expected positive result (is close to an expected test sample spectrum).
[0166] The method of training the machine learning algorithm may comprise allocating a classification to the reference spectra of the training data set. The classification may be binary (e.g. 0 or 1; pass / fail) or graded (e.g. a percentage score) other suitable classification. The classification may be carried out by visual inspection, or in any suitable manner. The classification may be allocated prior to training the algorithm.
[0167] The training data set may comprise between 50% and 99% of good spectra (a pass or 1), optionally between 55% and 95%, optionally between 60% and 90%, optionally between 70% and 90%, optionally between 80 % and 90 %. These examples are purely illustrative and any ratio of good to bad spectra may be used.
[0168] The method of training the machine learning algorithm may comprise identifying or selecting the at least one reference machine learning parameter.
[0169] The method of training the machine learning algorithm may comprise one or more pre-processing steps applied to some or all of the reference parameters of the reference spectrum or spectra, or to the reference spectrum / spectra.
[0170] The pre-processing step(s) may include removing one or more portions of a reference spectrum, which may be applied to one, some or all of the reference spectrum / spectra. The removal may be of one or two end portions of the spectrum to “cut” the spectrum to a desired range of frequencies.
[0171] The pre-processing step(s) may include applying a transformation or correction to a reference spectrum, which may be applied to one, some or all of the reference spectrum / spectra. The transformation or correction may be configured to correct the slope of the reference spectrum.
[0172] The pre-processing step(s) may include one or more steps of selectively modifying a reference parameter of a reference spectrum by applying a modification function thereto. The modification function may be configured to modify the reference parameter for one or more first conditions and to not modify the reference parameter for one or more second conditions. This step may be applied to some or all of the reference
[0173] 55763132-1spectra, and for some or all of the reference parameters thereof. In some examples, for a plurality of spectra, reference parameters are obtained for each spectra, and a modification function is applied to the obtained reference parameters. Depending on the original value of the reference parameter, modification to the value may occur. For example, one particular reference parameter may return a value of between -1 and 1, but it is known that the effect on the quality of the spectrum of values between -1 and 0, of this parameter, is substantially the same, and therefore it may be desirable for the machine learning algorithm to treat parameter values between -1 and 0 as the same as a value of 0. This avoids anomalously low values from adversely affecting the quality test.
[0174] The modification function may be a step function comprising the first condition and the second condition.
[0175] The modification function may be configured to: for obtained values of the reference parameter within the first condition, to attribute the same value or values for the modified reference parameter; and for obtained values within a second condition, to maintain the original value of the reference parameter.
[0176] The training may comprise applying a plurality of different modification functions to at least some of the reference spectra or to the reference parameter(s).
[0177] The pre-processing may comprise other functions, data processing, or the like, to modify the training data set of reference spectra prior to training the machine learning algorithm.
[0178] The training of the algorithm may comprise a first training step of the machine learning algorithm on a training portion of the training data set. The first training step may comprise one or more iterations using a training portion. Each iteration may include selecting a training portion. The training portion may be selected such that, for at least two iterations, the training portion is different. The selection of the training portion may be by random or substantially random selection from the training data set. There may be at least 2 iterations, or at least 10 iterations, or at least 30 iterations, or at least 50 iterations.
[0179] The first training step may comprise testing the machine learning algorithm on a test portion of the training data set. The testing may include having the machine learning algorithm make predictions on whether the reference spectra in the test portion are of sufficient quality. The step of testing may be carried out for at least some of the iterations of the first training step. The test portion, for each iteration, may exclude the training portion of that iteration.
[0180] 55763132-1The, or each training portion may comprise a plurality of the reference spectra of the training data set, optionally between 10% and 95%, optionally between 40% and 100%, optionally between 50% and 90%, optionally between 60% and 80%, optionally between 65% and 75%, optionally about 70%. The test data may make up some or all of the remaining training data set.
[0181] The, or each iteration of the first training step on a training portion may include a cross-validation (CV) process, which may be a nested CV, to train the algorithm on that training portion. The CV may be a two-fold, four-fold, five-fold, or any suitable number of folds. The CV process may include, for each fold, splitting the training portion into a subtraining portion and sub-test portion, and training the algorithm on the sub-training portion and testing the algorithm on the sub-testing portion.
[0182] The first training step may be an ensemble training method. The first training step may comprise bagging or bootstrap aggregation, in which the selection of a reference spectrum for inclusion in the training portion is done with replacement, such that the same reference spectrum may be selected for inclusion more than once. The bagging may be used for the CV step of selecting the, or each sub-training portion.
[0183] The machine learning algorithm may comprise one or more model hyperparameters. Training the machine learning algorithm may comprise setting and / or adjusting at least some of the one or more model hyper-parameters. The first training step may comprise setting and / or adjusting the one or more model hyper-parameters. The selection and / or adjustment of the model hyper-parameters may contribute to the training of the machine learning algorithm.
[0184] The step of adjusting the hyper-parameter(s) may be carried out after one or more iterations of the first training step. There may be a plurality of adjustments of the hyper-parameters carried out after an iteration of the first training step.
[0185] The method may comprise a second training step of training the machine learning algorithm. The second training step may be carried out after the first training step. The second training step may be carried out after the adjustment of the model hyperparameters.
[0186] The second training step may comprise training the machine learning algorithm on a training portion of the training data set. The second training step may comprise one or more iterations using a training portion. Each iteration may include selecting a training portion. The training portion may be selected such that, for at least two iterations, the training portion is different. The selection of the training portion may be by random or
[0187] 55763132-1substantially random selection from the training data set. There may be at least 2 iterations, or at least 10 iterations, or at least 30 iterations, or at least 50 iterations.
[0188] The second training step may comprise testing the machine learning algorithm on a test portion of the training data set. The testing may include having the machine learning algorithm make predictions on whether the reference spectra in the test portion are of sufficient quality. The step of testing may be carried out for at least some of the iterations of the second training step. The test portion, for each iteration, may exclude the training portion of that iteration.
[0189] The, or each training portion of the second training step may comprise a plurality of the reference spectra of the training data set, optionally between 10% and 95%, optionally between 40% and 100%, optionally between 50% and 90%, optionally between 60% and 80%, optionally between 65% and 75%, optionally about 70%. The test data may make up some or all of the remaining training data set.
[0190] The training portion of the training data set may be the same or different for the first and second training steps.
[0191] The, or each iteration of the second training step on a training portion may include a cross-validation (CV) process, which may be a nested CV, to train the algorithm on that training portion. The CV may be a two-fold, four-fold, five-fold, or any suitable number of folds. The CV process may include, for each fold, splitting the training portion into a sub-training portion and sub-test portion, and training the algorithm on the subtraining portion and testing the algorithm on the sub-testing portion.
[0192] The second training step may be an ensemble training method. The second training step may comprise bagging or bootstrap aggregation, in which the selection of a reference spectrum for inclusion in the training portion is done with replacement, such that the same reference spectrum may be selected for inclusion more than once. The bagging may be used for the CV step of selecting the, or each sub-training portion.
[0193] The method may comprise a third training step of the machine learning algorithm. The third training step may be carried out after the first training step and / or the second training step.
[0194] The third training step may comprise training the machine learning algorithm on a training portion of the training data set. The training portion of the third training step may be a larger training portion of the training data set than the first and / or second training steps, or all of the training data set. The third training step may comprise one or more iterations using a training portion. Each iteration may include selecting a training portion. The training portion may be selected such that, for at least two iterations, the
[0195] 55763132-1training portion is different. The selection of the training portion may be by random or substantially random selection from the training data set. There may be at least 2 iterations, or at least 10 iterations, or at least 30 iterations, or at least 50 iterations.
[0196] The third training step may comprise testing the machine learning algorithm on a test portion of the training data set. The test portion of the third training step may be some or all of the training data set. The testing may include having the machine learning algorithm make predictions on whether the reference spectra in the test portion are of sufficient quality. The step of testing may be carried out for at least some of the iterations of the third training step. The test portion, for each iteration, may exclude or include at least a part of the training portion of that iteration.
[0197] The, or each training portion may comprise a plurality of the reference spectra of the training data set, optionally up to 100%, optionally between 10% and 100%, optionally between 20% and 90%, optionally between 40% and 100%, optionally between 50% and 100%, optionally between 60% and 100%, optionally between 65% and 100%, optionally about 100%. The test data may make up some or all of the remaining training data set.
[0198] The, or each iteration of the third training step on a training portion may include a cross-validation (CV) process, which may be a nested CV, to train the algorithm on that training portion. The CV may be a two-fold, four-fold, five-fold, or any suitable number of folds. The CV process may include, for each fold, splitting the training portion into a sub-training portion and sub-test portion, and training the algorithm on the subtraining portion and testing the algorithm on the sub-testing portion.
[0199] The third training step may be an ensemble training method. The third training step may comprise bagging or bootstrap aggregation, in which the selection of a reference spectrum for inclusion in the training portion is done with replacement, such that the same reference spectrum may be selected for inclusion more than once. The bagging may be used for the CV step of selecting the, or each sub-training portion.
[0200] The third training step may comprise setting and / or adjusting at least some of the one or more model hyper-parameters. The step of adjusting the hyper-parameter(s) may be carried out after one or more iterations of the third training step. There may be a plurality of adjustments of the hyper-parameters carried out after an iteration of the third training step.
[0201] The hyper-parameter(s) determined or adjusted in the third training step may be the same or different to those tuned in the first training step. The hyper-parameter
[0202] 55763132-1corresponding to the number of features randomly chosen in each split of the CV may be adjusted in the third training step.
[0203] The method of training the machine learning algorithm may comprise determining a discriminatory threshold, which may be carried out at the third training step. This step may depend on receiver operating characteristic (ROC) data or curves.
[0204] The third training step may produce the machine learning algorithm used to determine the quality of the spectrum.
[0205] The first training step, the second training step and / or the third training step may include a decision tree method, which may be an ensemble of decision trees. The first, second and / or third training steps may include a supervised learning algorithm. The decision tree method may be a random forest model, or any suitable algorithm.
[0206] The method may comprise machine learning algorithms trained for different modes or samples.
[0207] The processor element may be operable to adjust the threshold at which a spectrum is deemed of acceptable quality. The processor element may be operable to adjust the quality scoring mechanism of the measured spectrum. The threshold for classification of the measured spectrum may be adjustable. The binary classification of the machine learning algorithm may be adjustable.
[0208] The method may comprise highlighting and / or removing measured spectrum / spectra that are not of acceptable quality. The removal may be by deletion, archiving or the like. The processor element or the machine learning model may carry out this step.
[0209] The numbering of the training steps is intended to differentiate the training steps and not to convey any order or preference unless the context explicitly provides otherwise. It is possible in some examples that other steps may be applied in addition to those described here. Furthermore, it is in no way essential that all of the training steps are required, although some embodiments may make use of all the training steps described here. It is also envisaged that the specific order of the steps may be different than as described here.
[0210] The machine learning algorithm, for a mode of operation, may be trained on reference spectrum / spectra corresponding to substantially the same type of sample. The machine learning algorithm, for the second mode of operation, may be trained on reference spectrum / spectra of the second sample. The machine learning algorithm, for the third mode of operation, may be trained on reference spectrum / spectra of the third sample.
[0211] 55763132-1The at least one reference machine learning parameter may be generated, at least in part, by the training of the machine learning algorithm. At least some of the reference machine learning parameters may be generated, at least in part, by the training of the machine learning algorithm. The at least one reference machine learning parameter may be generated by the training of the machine learning algorithm from one or more, or a plurality of reference spectra.
[0212] At least some of the reference spectra may be obtained from tests carried out using the spectrometer.
[0213] In some examples at least some of the reference spectra for use with a mode of operation may be obtained from a spectrometer carrying out that mode of operation. In some examples, at least some of the reference spectra for use with a sample may be obtained from a spectrometer on a same type of sample.
[0214] The training data set may comprise at least 100 reference spectra, or at least 1,000 or at least 10,000, or at least 15,000 reference spectra, or between 1 and 20,000 reference spectra, or between 1 ,000 and 20,000.
[0215] The reference spectra may include at least one of positive control data obtained from a positive control sample, negative control data obtained from a negative control sample, and clinical data, such as clinical blood, plasma or serum data. The reference spectra may include between 50% and 99% of clinical data, optionally between 55% and 95%, optionally between 60% and 90%, optionally between 70% and 90%, optionally between 75% and 85%.
[0216] The at least one spectrum parameter may have a real value. The at least one reference parameter may have a real value.
[0217] The step of determining the quality of the measured spectrum, when using machine learning, may comprise allocating a binary score (e.g. 0 or 1; or pass / fail) or a probability score (e.g. a percentage prediction as to the quality of the spectrum), or another suitable output, such as a plurality of scores (e.g. poor, acceptable, good).
[0218] The processor element may be configured to output the quality score. The processor element may be configured to output the quality score on the display element.
[0219] The processor element may be configured to output the quality score from the, or each mode of operation. The processor element may be configured to output the quality score from at least one of: the first, second and third modes.
[0220] At least some of the steps of the method of control testing may be computer implemented.
[0221] 55763132-1The steps of the method can be carried out in any order unless the context provides otherwise.
[0222] According to a second aspect of the disclosure, there is provided a method of quality testing a measured spectrum of a sample, the method comprising:
[0223] obtaining, by a processor element, a measured spectrum of a sample; identifying, by the processor element, at least one spectrum parameter of the measured spectrum;
[0224] comparing, by the processor element, the at least one spectrum parameter against at least one reference parameter;
[0225] determining, by the processor element, and based at least in part on the comparison of the at least one spectrum parameter to the at least one reference parameter, the quality of the measured spectrum.
[0226] According to a third aspect of the disclosure, there is provided a method of quality testing a measured spectrum of a sample, the method comprising:
[0227] obtaining, by a processor element, a measured spectrum of a sample; identifying, by the processor element, at least one spectrum parameter of the measured spectrum;
[0228] comparing, by the processor element, using a machine learning algorithm, the at least one spectrum parameter against at least one reference machine learning parameter;
[0229] determining, by the processor element, using the machine learning algorithm, and based at least in part on the comparison of the at least one spectrum parameter to the at least one reference machine learning parameter, the quality of the measured spectrum.
[0230] According to a fourth aspect of the disclosure, there is provided an apparatus for non-targeted analysis diagnostic measurement of a test sample, the apparatus comprising:
[0231] a spectrometer operable to receive a sample for measurement thereof;
[0232] a processor element configured, in use, to obtain a measured spectrum of the sample;
[0233] wherein the apparatus is operable to carry out:
[0234] (i) a first mode of control testing the spectrometer on a first, reference sample; and / or
[0235] (ii) a second mode of negative control testing on a second, negative control sample; and / or
[0236] 55763132-1(iii) a third mode of positive control testing on a third, positive control sample. According to a fifth aspect of the disclosure, there is provided an apparatus for testing the quality of a measured spectrum, the apparatus comprising:
[0237] a processor element;
[0238] wherein the processor element is operable to obtain a measured spectrum of a sample;
[0239] wherein the processor element is configured to identify at least one spectrum parameter of the measured spectrum;
[0240] wherein the processor element is configured to compare the at least one spectrum parameter against at least one reference parameter;
[0241] wherein the processor element is configured to determine, based at least in part on the comparison of the at least one spectrum parameter to the at least one reference parameter, the quality of the measured spectrum.
[0242] According to a sixth aspect of the disclosure, there is provided an apparatus for testing the quality of a measured spectrum, the apparatus comprising:
[0243] a processor element;
[0244] wherein the processor element is operable to obtain a measured spectrum of a sample;
[0245] wherein the processor element is configured to identify at least one spectrum parameter of the measured spectrum;
[0246] wherein the processor element is configured to compare, using a machine learning algorithm, the at least one spectrum parameter against at least one reference machine learning parameter;
[0247] wherein the processor element is configured to determine, using the machine learning algorithm, and based at least in part on the comparison of the at least one spectrum parameter to the at least one reference machine learning parameter, the quality of the measured spectrum.
[0248] According to a seventh aspect of the disclosure, there is provided an apparatus for testing the quality of a measured spectrum, the apparatus comprising:
[0249] a processor element configured to carry out at least one of the second or third aspects of the disclosure.
[0250] According to an eighth aspect of the disclosure, there is provided a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of at least one of the second or third aspects of the disclosure.
[0251] 55763132-1According to a ninth aspect of the disclosure, there is provided a machine learning algorithm configured, when run, to determine the quality of a measured spectrum.
[0252] According to a tenth aspect of the disclosure, there is provided a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of at least one of the second or third aspects of the disclosure.
[0253] According to an eleventh aspect of the disclosure, there is provided a spectrometer comprising the apparatus of the sixth or seventh aspects of the disclosure.
[0254] According to a twelfth aspect of the invention, there is provided a diagnostic method comprising:
[0255] running a method of control testing according to the first aspect of the disclosure; and
[0256] running a diagnostic test using the spectrometer.
[0257] According to a thirteenth aspect of the disclosure, there is provided a method of spectroscopy comprising obtaining a measured spectrum of a sample by a processor element.
[0258] The method may comprise providing a spectrometer. In some examples, the method may comprise obtaining the measured spectrum by the spectrometer. In some examples the method may comprise providing the measured spectrum to the processor element. For example, by retrieving the measured spectrum from a data source, such as a server, computing device, memory device, or the like.
[0259] The method may comprise providing a sample to the spectrometer.
[0260] The step of obtaining the measured spectrum may include obtaining a previously-generated measured spectrum or generating the measured spectrum by performing a measurement on the sample.
[0261] The above summary is intended to provide examples and be non-limiting. The disclosure includes one or more corresponding aspects, embodiments or features in isolation or in various combinations whether or not specifically stated (including claimed) in that combination or in isolation. It should be understood that features defined above in accordance with any aspect of the present disclosure or below relating to any specific embodiment of the disclosure may be utilized, either alone or in combination with any other defined feature, in any other aspect or embodiment or to form a further aspect or embodiment of the disclosure.
[0262] 55763132-1BRIEF DESCRIPTION OF DRAWINGS
[0263] These and other aspects of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, wherein:
[0264] Figure 1 depicts a spectrometer configured to run a control test of control samples, in accordance with embodiments of the invention;
[0265] Figure 2 depicts a control sample substrate for use with the spectrometer of Figure 1;
[0266] Figure 3 depicts example measured spectra, one for a positive control sample and one for a negative control sample;
[0267] Figure 4 depicts an example measured spectrum deemed to be of sufficient quality, and another example measured spectrum deemed to be of insufficient quality, obtained from the same diagnostic test sample, with the two spectra offset for ease of comparison; and
[0268] Figure 5 is a flow chart including running control test(s) of the spectrometer followed by clinical diagnostic tests using the spectrometer.
[0269] DETAILED DESCRIPTION OF DRAWINGS
[0270] Figure 1 shows a spectrometer 1 comprising a processor element 2, a display element 4, a receiving portion 6 for a sample, and a user interface 8.
[0271] Embodiments of the invention include a method of control testing for a nontargeted analysis spectroscopic measurement, the method comprising:
[0272] (i) carrying out a first mode of control testing the spectrometer 1, independent of a subsequent non-targeted analysis diagnostic measurement to be carried out on a test sample by the spectrometer 1, the first mode comprising:
[0273] providing a first, reference sample 10a to the spectrometer 1, wherein the first, reference sample 10a is configured to be independent of the test sample;
[0274] obtaining, by the processor element 2, a first measured spectrum of the first, reference sample 10a;
[0275] determining if the first measured spectrum 10a is of acceptable quality; and / or (ii) carrying out a second mode of negative control testing in respect of the subsequent diagnostic measurement to be carried out on the test sample by the spectrometer 1, the second mode comprising:
[0276] 55763132-1providing a second, negative control sample 10b to the spectrometer 1, wherein the second sample 10b is configured to produce a different measured spectrum to a measured spectrum of the test sample;
[0277] obtaining, by the processor element 2, a second measured spectrum of the second, negative control sample 10b;
[0278] determining if the second measured spectrum is of acceptable quality; and / or (iii) carrying out a third mode of positive control testing in respect of the subsequent diagnostic measurement, the third mode comprising:
[0279] providing a third, positive control sample 10c to the spectrometer 1, wherein the third sample 10c is configured to produce a same or similar measured spectrum to the measured spectrum of the test sample;
[0280] obtaining, by the processor element 2, a third measured spectrum of the third, positive control sample 10c;
[0281] determining if the third measured spectrum is of acceptable quality.
[0282] To avoid any confusion with the traditional meaning of positive and negative controls, the term “negative control” in this context means a control sample that will produce a spectrum that is deemed to be sufficiently different to an expected measured spectrum of the test sample. The term “positive control” in this context means a control sample that will produce a spectrum that is sufficiently close to or the same as an expected measured spectrum of the test sample. That is, the use of positive and negative controls does not imply any diagnostic result.
[0283] In some embodiments, the method comprises carrying out the subsequent non targeted analysis diagnostic measurement of the test sample. The non-targeted analysis is a non-specific spectroscopic measurement in which one or more or a plurality of unknown analytes are to be measured.
[0284] The processor element 2 is configured to determine the quality of the first measured spectrum, the second measured spectrum and / or the third measured spectrum, which will be described in more detail later. The / each step of determining the quality of the first measured spectrum, the second measured spectrum, and / or the third measured spectrum comprises:
[0285] identifying, by the processor element 2, at least one spectrum parameter of the measured spectrum; and
[0286] comparing, by the processor element 2, the at least one spectrum parameter against at least one reference parameter.
[0287] 55763132-1Determining the quality of the, or each measured spectrum, by the processor element 2, is based at least in part on the comparison of the at least one spectrum parameter to the at least one reference parameter. As will be described later, the first mode uses numerical comparison, and the second and third modes use machine learning for this part of the method.
[0288] The processor element 2 generates a quality score of the first measured spectrum and / or the second measured spectrum and / or the third measured spectrum, each of which is a discrete pass or fail value, but in other embodiments a scored range of values, such as a range of percentage values ora grading system such as poor, good, acceptable, excellent, or the like, may be used. It will be understood that other quality scores or indicators may be used. The quality score for the first mode is whether the first measured spectrum is sufficiently close to an expected measured spectrum for a reference control sample, and for the second and third mode, is whether a measured spectrum is sufficiently close to an expected measured spectrum of the test sample (for diagnostic measurements), with the second mode expected to return a “fail” and the third mode expected to return a “pass”, if the controls have been performed correctly.
[0289] Although the processor element 2 is shown in Fig. 1 as embedded within the spectrometer 1, and is connected to and integrated with the spectrometer 1, in other embodiments, the processor element 2 may be or may be part of a computing device, a spectrometer, a spectroscopy system, a measurement apparatus, a measurement apparatus comprising a spectrometer, a distributed computing system such as a cloud computing system, or the processor element 2 may be implemented virtually, such as using cloud computing or an emulator, or the processor element 2 may be remote from the spectrometer 1. It will be understood that in some embodiments of the method, the quality test may be carried out remotely by the processor element 2 obtaining measured spectra by data transfer from a spectrometer 1 or by other means, such as from a server, computing device, memory device, or the like.
[0290] In the embodiments of the method described here, the / each measured spectrum is obtained by the same spectrometer, with the control samples measured before diagnostic measurements are run on the test sample. In the embodiment of Fig. 1, the spectrometer 1 is an Attenuated Total Reflectance-Fourier Transform Infrared (ATR-FTIR) spectrometer 1, but other spectroscopy techniques may be employed and the invention is by no means limited to the use of ATR-FTIR. Other spectroscopic techniques include any infrared (IR) spectroscopy, vibration-based spectroscopy (e.g. Raman spectroscopy) or any spectroscopy that uses electromagnetic radiation to produce a
[0291] 55763132-1measured spectrum of a sample, or the like. The non-targeted analysis of the test sample is by ATR-FTIR, but in some other embodiments, targeted analysis and / or another form of spectrometer may be employed. In some embodiments, data may be obtained from a spectrometer for analysis (i.e. some embodiments do not require a spectrometer, as some embodiments are directed to methods of quality testing or analysis, computer programs, etc.).
[0292] The spectrometer 1 is configured for carrying out non-targeted analysis diagnostic measurements, and some embodiments of the invention include such diagnostic methods and hardware, and some embodiments include control testing of measured spectrum / spectra from at least one of the first to third samples and then carrying out diagnostic method(s) on the test sample.
[0293] In the embodiments illustrated and described here, the control tests of the invention are to determine if such diagnostic measurements will be, or are likely to be valid. In the embodiment of Fig. 1 , the spectrometer 1 is configured to carry out diagnostic measurements using one or more artificial intelligence (Al) or machine learning algorithms or processes to analyse the measured spectrum of the test sample to measure one or more analytes and to diagnose or infer one or more health conditions, such as cancer or the like, present in human or animal biological material, which may be blood, plasma, or serum, or other biological material. The diagnostic machine learning or Al is different from the quality testing machine learning, which will be described in more detail further below.
[0294] Embodiments of the method include a numerical comparative aspect, which in the embodiment of Fig. 1 is carried out without machine learning for the first mode, and a predictive quality test of the measured spectrum using machine learning for the second mode and the third mode, both of which are used to determine if a subsequent diagnostic measurement will be or is likely to be valid. It will be readily understood from a reading of this specification that some embodiments may only use the numerical comparison, some embodiments may only use the machine learning comparison, and some embodiments may use a combination thereof, as required. Likewise, some embodiments of the hardware may be configured to run one or both types of quality test.
[0295] For the control test and the subsequent diagnostic measurement(s), the spectrometer 1 measures absorbance versus frequency (in wavenumbers), of the sample 10. In at least some embodiments, the spectrometer obtains a measured spectrum including frequencies corresponding with at least some or all molecules of interest contained within human / animal blood, plasma, or serum so that the spectrum
[0296] 55763132-1contains a signal indicative of proteins (e.g., albumin, globulin), lipids (e.g., cholesterol), carbohydrates (e.g., glucose), deoxyribonucleic acid (DNA), ribonucleic acid (RNA), sugars, drugs, metabolites, phosphate, and / or any other component in the human / animal blood, plasma, or serum. It will be appreciated that this is illustrative and other embodiments can use different spectral parameters, such as other intensity values, different ways of expressing the frequency, and be configured to produce different ranges of spectra, and it is not essential to target all the molecules of interest in blood, plasma or serum, as there are many applications where the control testing and quality measurement of a spectrum may be required.
[0297] In the embodiment of Fig. 1, a sample 10 is provided to the sample receiving portion 6 for each repetition, or each mode, of the method. Some embodiments of the invention do not require this step as the spectrum / spectra will be obtained remotely.
[0298] The processor element 2 and the spectrometer 1 of Fig. 1 are configured to automatically run the step of obtaining and determining the quality of the first, second and / or third measured spectrum in response to a first user input step provided to the processor element 2 via the user interface 8. In some embodiments, the processor element 2 runs the / each quality check without the participation of the spectrometer 1 , by obtaining or being provided with a measured spectrum to check the quality thereof.
[0299] The processor element 2 and the spectrometer 1 form a system configured to run the method of control testing on the control samples and to determine the quality of measured spectra thereof, and to carry out a method of diagnosis on the test sample, so that the validity of the diagnostic method can be predicted before it is run on the test sample. This is advantageous as the control samples do not use any of the test samples (e.g. do not comprise patient samples) and the predicted validity can help to mitigate failed diagnostic measurements or test sample (e.g. patient sample) waste. In some known techniques, pooled patient samples may be used for controls, but gathering such pooled samples may not be practical, whereas in some embodiments, the controls are easily obtainable, repeatable, and consistent, among other advantages. It will be appreciated that the controls described here can be prepared in advance of a diagnostic measurement, as they do not rely on a delivery of test samples. Furthermore, to prepare or select the control samples only requires knowledge of the carrier of the test sample (e.g. blood, plasma, serum, or other carrier), and for the reference sample, no biological material is used at all. The controls are also not specific to the condition that is to be diagnosed - the controls are instead focussed on determining if the spectrometer is working correctly (first mode), the operator can prepare a true negative result that will be
[0300] 55763132-1deemed to not sufficiently match an expected measured spectrum of the test sample (second mode) and the operator can prepare a sample and perform a measurement of a true positive result that will be deemed to be close to an expected measured spectrum of the test sample.
[0301] The display element 4 is common to the spectrometer 1 and the processor element 2 in Fig. 1. Other arrangements are possible than depicted in Fig. 1.
[0302] The first user input step starts each mode and begins the step of obtaining a measured spectrum for each mode by the spectrometer 1. The determination of the quality of the measured spectrum for each mode is automated by the processor element 2 without any further user input, from the first user input to the determination of the quality of the measured spectrum. In other embodiments, there may be other user input steps.
[0303] The control testing method can be carried out again to obtain further spectra from a sample and to determine the quality of each measured spectrum thereby obtained. The or each of the first, second and third modes of the method are repeatable each time by a further user input step due to the automation described above. Fig. 4 shows such a repeat for the sample, in the third mode, in which one run has passed and one run has failed the quality check. It will be appreciated that a control testing protocol may deem a number of runs of each mode to be desirable to then either proceed to diagnostic measurements, or to perform troubleshooting, or to take another action. In the embodiment of Fig. 1, the processor element 2 and the spectrometer 1 are configured to prevent a diagnostic measurement of the test sample unless the first measured spectrum is of acceptable quality, and unless the second measured spectrum is of acceptable quality, and unless the third measured spectrum is of acceptable quality. The processor element 2 and the spectrometer 1 are configured to prompt a repeat of the first mode and / or the second mode and / or the third mode, as required, if at least one mode returns a spectrum of unacceptable quality, before permitting a diagnostic measurement of the test sample to be carried out. In some embodiments, it may not be required to carry out all three modes, and it may not be required to obtain a pass for all three modes to carry out diagnostic measurements.
[0304] The user interface 8 may take any suitable form, including button(s), a touch screen interface, keyboard, mouse, or the like.
[0305] In the embodiments illustrated and described here, the method of quality testing comprises a first mode, a second mode and a third mode. It is not essential to run all three modes, and in some embodiments, only one or some of the modes may be run
[0306] 55763132-1and the hardware may be configured or operable to run one or more, some or all of the first, second and third modes.
[0307] In the embodiment of Figs. 1 and 2, the method comprises running the first mode, the second mode and the third mode on a plurality of samples 10a, 10b, 10c which are located on a common substrate 12 (a silicon-based substrate including a housing and configured for use with the ART-FTIR spectrometer 1), which is one sample 10a, 10b, 10c for each mode. In other embodiments, the samples may be separate. The spectrometer 1 is configured to generate multiple spectra from a plurality of samples 10a, 10b, 10c, which may be one spectrum per sample 10a, 10b, 10c, or a plurality of spectra from each sample 10a, 10b, 10c, as required.
[0308] The processor element 2 can output the quality score from each mode of operation in any suitable manner, such as by displaying it on the display element 4, saving and / or transmitting the data to a suitable location.
[0309] The first mode is a method of quality testing the spectrometer 1 using the first, reference sample 10a, which is a blank or neutral control sample. The second mode is a method of quality testing using the second, negative control sample 10b. The third mode is a method of quality testing using the third, positive control sample 10c. The method may comprise other modes than those listed here and may use different control samples than those described here.
[0310] In the embodiment of Fig. 1 , the first mode is a hardware and software verification mode for determining if the spectrometer 1 is working as expected independent of a biological sample, which is a test that is independent of biological material, reagents, or the like. In some embodiments, this mode may include verifying the operation of any other hardware / software parts of the system in addition to the spectrometer 1. The first mode is independent of a subsequent non-targeted analysis diagnostic measurement to be carried out on the test sample by the spectrometer 1 and the first mode is independent of the second mode and / or the third mode. In this example, the performance of the first mode, including the selection of the properties of the first, reference sample does not depend on the second and / or third modes. In some examples, the first mode and the first reference sample are independent of the second mode, the third mode, and the subsequent measurement to be carried out on the test sample.
[0311] In the embodiment of Fig. 1 , the spectrometer 1 and the processor element 2 are configured to not permit the second mode and / or the third mode to be carried out unless the first mode returns an acceptable first measured spectrum, because if the spectrometer and / or software thereof is not functioning correctly, this saves the
[0312] 55763132-1preparation and measurement of the second and third control samples, which generally speaking take more time and effort to prepare than the first, reference sample, as the first sample will typically be a stable reference that does not need to be prepared each time.
[0313] The second mode is a control test for determining if the negative control sample 10b will be identified as such by the spectrometer 1 because the measured spectrum is not close enough to an expected measured spectrum of the test sample. The spectrometer 1 and the processor element 2 are configured to provide a fail result for a successful run of the second mode because the second measured spectrum is sufficiently different from an expected spectrum of the test sample. The spectrometer 1 and the processor element 2 are configured to provide a pass result if the second measured spectrum is sufficiently similar to an expected spectrum of the test sample. The second mode is for determining if a user can prepare a true negative sample that will not be identified by the spectrometer 1 as being comparable to an expected measured spectrum of the test sample, and a successful result is a fail being obtained which indicates correct preparation of a negative control.
[0314] The third mode is a quality test for determining if a positive control sample will be identified as such by the spectrometer 1 because the measured spectrum is close enough to an expected measured spectrum of the test sample. The spectrometer 1 and the processor element 2 are configured to provide a pass result if the third measured spectrum is sufficiently similar to an expected measured spectrum of the test sample. The third mode is for determining if a user can prepare a true positive sample and is operating the spectrometer 1 correctly to enable a measured spectrum to provide an expected positive match result for the positive control sample 10c, such that it can be expected that the measurement of the test sample will be performed correctly.
[0315] Typically, the quality testing method will comprise running the first mode then the second mode then the third mode, as if the hardware is not working as expected, there may be no value in running the second and third modes. Each mode produces a quality result of the measured spectrum of that mode, and if the spectra are deemed to pass, a diagnostic method can be run (or other method) from the spectrometer with a degree of confidence that the result will be valid. In at least some embodiments, it may not be necessary for at least one of the modes to have the processor element 2 carrying out the quality check, as it may be sufficient for visual inspection to be used to determine the quality of a measured spectrum for at least one mode. The term “deemed to pass” means provide an expected result, which is a fail for the negative control.
[0316] 55763132-1The control testing method may comprise obtaining more than one spectrum for each sample 10a, 10b, 10c, as it may be desirable to use averaging when determining the quality of a measured spectrum of a sample.
[0317] Fig. 3 shows measured spectra, one obtained from the second mode (negative control; labelled BSA) and one from the third mode (positive control; labelled LHPS), each comprising a plurality of data points of absorbance against frequency in wavenumbers, between about 1,000 cm-1and about 3,500 cm-1. The BSA trace is a “fail” because it is not close to an expected measured spectrum of the test sample, and the LHPS trace is a “pass” as it is close to the expected measured spectrum. It will be described in detail later how this is determined.
[0318] Fig. 4 shows example spectra from the same diagnostic test sample, except one spectrum has passed (CQT pass) and one spectrum has failed the quality test (CQT fail). It is clear from Figure 4 that the spectrum which has failed is much noisier around the 3500-2800 cm-1region when compared to the spectrum which has passed. In some examples, this region contains biologically relevant peaks important to the diagnostic model which cannot be remedied using spectral pre-processing. This therefore demonstrates the successful use of the test in identifying spectra from a test sample which require repeat analysis or in which perhaps further control testing is required.
[0319] The samples 10a, 10b, 10c for the quality test are control samples and the diagnostic test sample is provided to the spectrometer 1 after the method of control testing. The diagnostic test sample comprises biological material, such as human / animal blood, plasma or serum, or other biological material, although this is an example, and other diagnostic, or other, samples may be tested once the quality test has been run.
[0320] In some embodiments, a plurality of or three or more control samples may be used for the or each mode. Each sample may be measured more than once, as required.
[0321] The first, reference sample 10a comprises no subject test samples, such as diagnostic test samples. A subject test sample means a sample from a human or animal for whom a test result is to be obtained from the test sample. The second, negative control sample 10b comprises no subject test samples or portion of subject test samples. The third, positive control sample comprises no subject test samples or portion of subject test samples. The first, reference sample 10a, the second negative control sample 10b and the third, positive control sample 10c are devoid of subject test material intended for use with the test sample.
[0322] The first, reference sample 10a is a non-biological blank or neutral control sample. The first, reference sample 10a is substantially devoid of any biological material
[0323] 55763132-1present in the second, negative control sample 10b and the third, positive control sample 10c. The first, reference sample 10a is independent from and substantially different to the test sample. The first, reference sample 10a is substantially devoid of any biological material present in the test sample. The first, reference sample 10a is configured to be substantially biologically inert. The first sample 10a is selected or prepared to be independent of the second sample 10b and the third sample 10c material properties. In the embodiment of Figs. 1 and 2, the first sample 10a is epoxy resin on the substrate 12, but in other examples may comprise or may be a plastics or polymer-based sample, or may comprise or may be a plastic or polymer sample. The first sample 10a is a solid sample, but in other embodiments may be a solid and / or liquid sample. The first sample 10a is a synthetic sample.
[0324] The first, reference sample 10a is configured to be re-usable within 2 years of its preparation, and is stable within this time period. The first mode can therefore generally be carried out by re-using the same first sample 10a with minimal preparation over a long period of time, which is advantageous over other forms of reference sample. Many spectra can be obtained from the first reference sample 10a over time, and then when carrying out the first mode, the quality check is whether the first measured spectrum is as expected, by numerical comparison in some embodiments, or it may be sufficient in some examples to use visual inspection.
[0325] The epoxy resin first, reference sample 10a is configured to produce a plurality of sharp spectral peaks when measured by the spectrometer 1 and at least one signal-free region. Any suitable first, reference sample that produces a signal peak region and a signal-free region for noise analysis will generally provide a sample that is suitable for carrying out the first mode, and the use of the term “sharp peaks” will be understood by those of skill in the art to mean discernible signal peaks suitable for comparing one spectrum to another spectrum.
[0326] The second sample 10b is a liquid bovine serum albumin (BSA) solution and therefore contains animal biological material. The second sample 10b is different to the test sample. The second sample 10b comprise a comparable component (albumin) to the third sample and to the test sample (which are human serum samples). Albumin is the largest fractional component of human serum and therefore makes a significant contribution to the resultant IR spectrum. It is advantageous for the spectrum to be similar but far enough removed from the expected spectrum of the third sample 10c and the test sample, to provide a working test for the second mode. If the expected spectrum of the second mode was too far from the expected spectrum of the third sample and the test
[0327] 55763132-1sample, it would not provide as much of a useful test of the preparation and measurement of a true negative sample, and likewise if the expected spectrum of the second sample 10b was too close to the third sample and test sample, it may not be possible to differentiate the measured second spectrum from the expected spectrum of the third sample 10c and the test sample.
[0328] In other examples, other components or comparable components may be selected based on the third sample 10c and / or the test sample, which may include protein(s), an animal component comparable to a human component, at least a portion, or a comparable portion, of the third sample 10c and / or the test sample. The component or comparable component may include a serum component when the third sample 10c and test sample are serum samples, a plasma component when plasma is used, or a blood component when blood is used. The method of control testing therefore comprises selecting or preparing the second, negative control sample 10b based, at least in part, on one or more properties of the test sample. The component or comparable component may include human and / or animal component(s). The second sample 10b may comprise or may be blood, plasma, serum, bovine serum albumin (BSA), a BSA solution, bovine plasma albumin, a bovine plasma albumin solution, or any suitable biological material. In some embodiments, the second sample 10b can be a liquid and / or a solid sample.
[0329] The BSA solution of the second sample 10b is procured in lyophilised format. BSA is supplied to a user in a pre-weighed vial containing 10 mg of powdered BSA, which is reconstituted in 100 pL of distilled water by the user to create a 100 mg / mL solution. Reconstituted vials of BSA remain stable in a refrigerated environment (+2 to +8 °C) for four weeks, corresponding with the recommended usage period. It will be understood that other BSA solutions may be used. The second sample 10b is therefore stable for about 1 month.
[0330] The third sample 10c is a biological sample comprising a lyophilised human pooled serum (LHPS) solution. This is because the test sample is a serum sample. In other embodiments, the third sample 10c could comprise human or animal biological material, which may include blood, plasma, serum, lyophilised human pooled serum (LHPS), a LHPS solution, lyophilised human pooled plasma (or solution thereof) or any suitable biological material.
[0331] The third sample 10c is different to the test sample. The third sample 10c comprises substantially all the carrier components of the test sample, because both are serum samples (serum being the carrier), and thus the spectrum of the third sample 10c will be close to identical to the test sample (i.e. close to standard patient serum), but in
[0332] 55763132-1other embodiments the third sample 10c may comprise one or more or a plurality of carrier components of the test sample. The shared components of the third sample 10c and the test sample are biological components, which are components of serum, but other examples may use blood, or plasma, or other shared components. The use of “shared” does not mean the use of a common sample - it means the use of the same type of component. For example, using two different serum samples will mean that substantially the same shared components are present, although there will of course be some difference between the samples, but the shared components mean that the spectra will be sufficiently similar to enable a positive control to be carried out.
[0333] The third sample 10c and the test sample are human serum samples, but the third sample 10c does not come from the test samples to be measured, although in some examples it is possible to use the test samples in this way, it is advantageous to not have to rely on the use of test sample material, particularly in situations where a limited amount of test material may be available.
[0334] The third sample 10c is generally configured to match the carrier type of the test sample (e.g. serum, plasma, blood, or other carrier). In some examples, depending on the carrier, the third sample 10c may have at least one more component in common, or comparable component in common, with the test sample than the second sample 10b, or 5 or more, 10 or more, 20 or more, or any suitable number of more common components or comparable components. That is, the third sample 10c is chosen to provide a measured spectrum that will be sufficiently closer to the expected measured spectrum of the test sample than the closeness of the second measured spectrum to the expected spectrum of the test sample.
[0335] The LHPS is obtained in powder format, and supplied to the user in a preweighed vial containing 5 mg of LHPS, which is reconstituted in 100 pL of distilled water by the user to create a 50 mg / mL solution. Reconstituted vials of LHPS remain stable in a refrigerated environment (+2 to +8 °C) for four weeks, corresponding with the recommended usage period. The third sample 10c is therefore stable for about 1 month. It will be understood that other solutions or samples may be used.
[0336] The third sample 10c is a liquid sample, but in other examples may be a liquid and / or a solid sample.
[0337] The test sample is a liquid human serum sample, but in other embodiments the test sample may be different, and the second and / or third samples 10b, 10c would then be different (e.g. switching to plasma may use bovine plasma albumin and lyophilised human plasma solution).
[0338] 55763132-1The test sample has substantially the same volume as the third sample 10c and the second sample 10b.
[0339] Embodiments of the method of control testing comprise selecting or preparing the third, positive control sample 10c based, at least in part, on one or more properties of the test sample.
[0340] The first and second samples 10a, 10b are devoid of reagents.
[0341] Advantageously, the first, second and third samples 10a-10c are controls for quality control that are consistent, stable and readily available.
[0342] The at least one reference parameter is used in the first mode, each of which is a reference spectral parameter derived from, at least in part, a plurality of reference spectra. For the first mode, the step of comparing the at least one measured spectrum parameter against the at least one reference parameter includes comparing 23 measured spectrum parameters to 23 reference parameters (comparing each measured spectrum parameter against one reference parameter). The first mode uses a numerical comparison in this respect, and does not use a machine learning algorithm to determine the quality of the measured spectrum. Instead, the processor element compares each measured parameter in turn against the corresponding reference parameter, and there is a threshold value or range whereby the measured parameter will be deemed to pass or fail. If all parameters pass, the measured spectrum passes, whereas if one measured parameter is a fail, the spectrum will fail. In other embodiments, this specific method of determining the quality of a measured spectrum may be configured differently.
[0343] The reference parameters used for the first mode are each numerical values or numerical ranges, which may be open at one end, or a closed range.
[0344] In the first mode of the embodiment of Fig. 1 , the reference parameters are listed in the table below, with the first two rows also including the comparison associated with the reference parameter. The upper and lower bounds provide the test for whether a measured parameter is of sufficient quality or not.
[0345]
[0346] 55763132-1
[0347]
[0348] The reference parameters can include one or more of the following:
[0349] the intensity values for a range of frequencies;
[0350] an intensity value for a frequency;
[0351] a location of one or more expected peak values or trough values;
[0352] a peak intensity value or trough intensity value (maximum value or minimum value) at one or more frequencies;
[0353] a prominence of one or more peak intensity values or trough intensity values for a range of frequencies;
[0354] a signal to noise ratio of an intensity value, or intensity values of the spectrum, which may include the signal to noise ratio of a peak intensity value or trough intensity value against another part of the spectrum, such as a non-signal part or non-peak / trough part of the spectrum; and / or
[0355] a relative peak to trough value within a range of frequencies.
[0356] In some examples, including the embodiment of Fig. 1 , the reference parameters can include multiple parameters of the same type. For example, intensity values for different ranges of frequencies, locations of peak values, etc.
[0357] In the embodiment of Fig. 1 and with reference to the above table, a prominence is a vertical distance between the peak intensity value and its lowest contour line, or
[0358] 55763132-1between a trough intensity value and its highest contour line. It will be understood that the prominence can be defined in many different ways.
[0359] Each identified spectrum parameter of the measured spectrum is selected by the processor element 2 to correspond substantially to a reference parameter.
[0360] The identified spectrum parameters are identified as contributing to the quality of the measured spectrum. The processor element 2 is configured to do this by being programmed to select the spectrum parameters (for all three modes) but in other embodiments this may involve some manual input to the processor element 2 if required.
[0361] The identified spectrum parameters are identified based on the mode of operation and the sample that is to be, or has been, used to obtain the measured spectrum. In some embodiments, the method comprises pre-configuring or selecting information about the mode of operation, the test to be run, and / or the sample, in order to set up the processor element 2 for carrying out the quality test method.
[0362] The reference parameters used for the first mode are different to those used for the second mode and the third mode. In other examples, the same reference parameters might be used between different modes.
[0363] The first mode comprises numerically comparing the 23 identified spectrum parameters to the 23 reference parameters to generate a binary quality measure (pass or fail), but in other embodiments, the first mode may generate a probabilistic measure of the quality, or any suitable quality measurement.
[0364] The numerical comparison of the first mode comprises arithmetic and vector operations, such as comparing the measured parameter numerically to a range of acceptable values, or a vector operation such as obtaining an Euclidean norm of the difference of a measured spectrum parameter to a reference spectrum.
[0365] The performance of the first mode was assessed in two different ways. The first involved a set of 60 good quality epoxy spectra which were used to assess the sensitivity of the test to identify good quality spectra. The second involved a set of 900 control spectra (LHPS and BSA) which were used as a check to assess the robustness of the first mode test against different types of data.
[0366] The epoxy test data that were used for testing the first mode comprised 60 spectra measured from 60 slides (1 spectrum per slide). The quality of these spectra was assessed by-eye by a laboratory technician after measurement and was found to be sufficient. Thus, for the purpose of testing the first mode, the quality labels for all spectra were assumed to be a pass (good quality). Epoxy spectra were pre-processed by cutting
[0367] 55763132-1the spectrum to 1000 cm-1 to 3000 cm-1 and correcting the slope thereof using an appropriate epoxy reference spectrum.
[0368] The control set that was used for testing the first mode comprised 900 spectra (450 LHPS and 450 BSA). The LHPS samples were prepared at 50 mg / mL concentrations, while the BSA samples were prepared at 100 mg / mL. Within the scope of testing the first mode, their quality labels were assumed to be ‘FAIL’ because clearly they would not be expected to produce the same outcome as the first, epoxy-based reference sample.
[0369] When analysing these spectra using the first mode, 100% of the epoxy spectra passed and were deemed good quality, while 100% of the control spectra (BSA and LHPS) failed and were deemed bad quality. The latter provided a check that the test fails spectra which are different to epoxy reference spectra.
[0370] This method of checking the first mode is an example, and there are other ways of testing the first mode. Likewise, the use of an epoxy sample as the first sample is one of many neutral or blank control samples that can be employed in the first mode of the quality testing method.
[0371] In the second and third modes, the steps of comparing the at least one spectrum parameter against the at least one reference parameter; and determining the quality of the measured spectrum are carried out using a quality testing machine learning algorithm (different to a diagnostic machine learning or artificial intelligence algorithm that may be used for diagnostic analysis or measurement of the test sample). Therefore, in the embodiment of Fig. 1 , the quality testing method comprises carrying out of the first mode of the method using a numerical comparison method and carrying out the second and third modes of the method using the quality testing machine learning algorithm.
[0372] In the embodiment of Fig. 1, the processor element 2 comprises the quality testing machine learning algorithm for determining if a measured spectrum from a sample is of acceptable quality, and is operable to run the quality testing machine learning algorithm. In other embodiments, the processor element 2 could run a machine learning algorithm on a further device, such as a further computing device. The machine learning algorithm may be run in various ways, such as on a computing device that the processor element 2 is part of, on another computing device, on an emulator or cloud server or the like, and in some embodiments, different functions or parts of the machine learning algorithm may be run on different devices. It will also be understood that the term machine learning algorithm does not exclude artificial intelligence, deep learning, or other techniques applicable to the present invention. It is also possible that more than
[0373] 55763132-1one algorithm may be used. It is also envisaged that in some examples multiple runs of the machine learning algorithm may be carried out to obtain an average result or to allow for erroneous runs to be discarded.
[0374] In the embodiment of Fig. 1, the quality testing machine learning algorithm is a random forest model (an example of a supervised learning, decision tree method, using an ensemble of decision trees). However, the use of other machine learning techniques and models is also possible, and this specific example is provided for illustrative purposes to understand how the second and third modes may be carried out.
[0375] The quality testing machine learning algorithm has been trained for each of the second and third modes of operation of the quality method, and is configured for use with the specific second and third samples 10b, 10c. The quality testing machine learning algorithm can be trained on other sample types and modes of operation. In some examples, the carrying out of the quality testing method may include user input to inform the processor element 2 of which mode or sample is being carried out / tested, or this may be automatically detected by the machine learning algorithm due to the similarity of the measured spectrum to an expected result for a mode or sample.
[0376] In the second and third modes, the reference parameters are reference machine learning parameters. In some embodiments, these may be the same parameters as the first mode, or any of the example reference parameters.
[0377] The embodiment of Fig. 1 comprises using the first mode using the reference parameters, and the second mode and the third mode using machine learning and the reference machine learning parameters. The first mode is a numerical comparison test and the second mode is a machine learning prediction or classification test.
[0378] Each reference machine learning parameter includes one or more spectral parameters derived from, at least in part, a plurality of reference spectra.
[0379] Each reference machine learning parameter is a numerical value or a numerical range. The numerical range may be open at one end, or be a closed range.
[0380] In the second and third modes, 135 spectrum parameters are compared against 135 reference machine learning parameters. The quality testing machine learning algorithm uses the comparison of measured parameters to reference parameters to determine the quality of the measured spectrum or spectra. In the embodiment of Fig. 1 , the reference machine learning parameters are detailed in the table below.
[0381]
[0382] 55763132-1
[0383]
[0384] 55763132-1
[0385]
[0386] 55763132-1
[0387]
[0388] 55763132-1
[0389]
[0390] It will be understood that the comparison of the measured parameters to the reference machine learning parameters may be carried out in many ways, and it may not involve a direct comparison, but the machine learning algorithm may use the reference machine learning parameters to learn how to determine the quality of a measured spectrum.
[0391] It will be understood that the machine learning algorithm may be configured to determine the quality of a measured spectrum without directly looking up a reference parameter, but in this example, the comparing of the at least one spectrum parameter against at least one reference machine learning parameter still effectively takes place as the machine learning algorithm will need to make decisions based on reference parameters.
[0392] 55763132-1In other embodiments, the at least one reference machine learning parameter can include one or more of:
[0393] a location of a peak intensity value or trough intensity value within a range of frequencies;
[0394] a location of a maximum or minimum intensity value within a range of frequencies;
[0395] a value of a peak or trough for a range of frequencies;
[0396] a location of a peak intensity value or trough intensity value within a region of the spectrum associated with a substance (e.g. within an absorbance band of the substance);
[0397] a location of a local or global peak or global trough within a range of frequencies; the intensity value of the local or global peak / trough;
[0398] spectral intensity values for a range of frequencies;
[0399] a transformation parameter used to apply a transformation or correction to the measured spectrum or spectra, or to the measured spectrum parameter(s);
[0400] a ratio of the intensity value of a local maximum within a range of frequencies to a local minimum within a range of frequencies, wherein the ranges may be different or substantially the same;
[0401] a signal to noise ratio, which may include the signal to noise ratio of a peak intensity value or trough intensity value against another part of the spectrum, such as a non-signal part or non-peak / trough part of the spectrum;
[0402] a prominence of one or more peak values or trough values;
[0403] a sum of prominences of a local minima and maxima within a range of frequencies;
[0404] a width of a peak or trough for a range of frequencies, which may be for a substance band.
[0405] In other embodiments, the reference machine learning parameter(s) may include any features or options of the reference parameters and vice versa.
[0406] Any suitable range of frequencies may be selected to correspond with an absorbance band of a substance.
[0407] For the signal to noise ratio, the signal portion may include an absorbance band of a substance or a range of values for a range of frequencies, and the noise portion may be outside of the absorbance band, or at a non-peak / non-trough section, or at a nonsignal section.
[0408] 55763132-1The transformation parameter may be one or more parameters of a function for correcting a slope of the measured spectrum. In this example, the transformation applied to the measured spectrum may be compared with the reference value or range of values that are deemed acceptable for this parameter.
[0409] The substance region or band may be any suitable substance band used with spectroscopic analysis, such as an amide band, such as amide I, or amide II. These are examples and are not intended as being essential to the invention.
[0410] In the second and third modes, the comparison used by the algorithm of the reference machine learning parameters is included in the table. When using the quality testing machine learning algorithm, the step of comparing the spectrum parameters to the reference machine learning parameters can comprise measuring one or more of: a peak or trough x-shift (frequency shift) of the measured spectrum parameter against the reference machine learning parameter;
[0411] a peak or trough y-shift (difference in intensity) of the measured spectrum parameter against the reference machine learning parameter;
[0412] a numerical comparison or arithmetic operation between the measured spectrum parameter and the reference machine learning parameter;
[0413] a calculation of the degree of similarity to expected value(s), which may be a percentage value or the like;
[0414] comparison to mean values, which may be a mean reference parameter obtained from a plurality of reference parameters, where the mean may be for a range of frequencies;
[0415] an area difference to the mean, which may be for a range of frequencies;
[0416] a cosine distance to the mean, which may be for a range of frequencies;
[0417] a 1-norm distance to the mean, which may be for a range of frequencies; the number of intensity values that are outside of n standard deviations from the mean, where n can be e.g. 1, 2, 3, or any suitable value;
[0418] a ratiometric difference between a local maximum and a local minimum with a range of frequencies;
[0419] a sum of two or more prominences of local minimum / minima and local maximum / maxima in a range of frequencies;
[0420] determination if a parameter is above or below a threshold value, such as assigning a “1” if the width of an intensity peak is below a threshold width, and a “0” if above the threshold width.
[0421] 55763132-1In the embodiment of Fig. 1 , the mean is a mean value of a plurality of reference spectra, which effectively creates a mean reference spectrum for identifying the similarity of a measured spectrum to a mean reference spectrum. As detailed in the table above, the mean may be for a range of frequencies, such as a part of the mean reference spectrum.
[0422] For the second and third modes, each identified spectrum parameter of the measured spectrum is selected by the processor element 2 to correspond substantially to at least one reference machine learning parameter.
[0423] In some embodiments, at least one, or some, or all of the reference machine learning parameters may be different between the second mode and the third mode.
[0424] The second mode and / or the third mode each comprise comparing the identified spectrum parameters to the reference machine learning parameters to determine the quality score, which is a discrete pass or fail value as to whether a measured spectrum is sufficiently close to an expected spectrum of the test sample, but other ways of expressing the quality may be employed, such as another binary measure, or a scored range of values, such as a range of percentage values ora grading system such as poor, good, acceptable, excellent, or the like. It will be understood that other quality scores or indicators may be used.
[0425] For the second and third modes, the spectrum parameters of the measured spectrum are selected to correspond to the reference machine learning parameters.
[0426] A description of how the quality testing machine learning algorithm of the embodiment of Fig. 1 was trained will now be provided.
[0427] The machine learning algorithm is trained using a plurality of reference spectra, which form a training data set.
[0428] The machine learning algorithm is trained for determining the quality of different expected results, namely does a positive control result in a spectrum close to an expected test sample spectrum and does a negative control result in a spectrum that is not close to an expected test sample spectrum.
[0429] The reference spectra of the training data set included clinical data and positive and negative control data (using LHPS and BSA respectively). Each patient sample was prepared by pipetting 3 pL of serum onto three sample wells of a sample slide. Prepared slides were placed in a drying unit incubator (Thermo Scientific™ Heratherm™, USA) at 35 °C for one hour, to control the dehydration process of the serum droplets. The dried sample slides were loaded into the sample receiving portion 6 of the spectrometer 1 and prepared for spectral collection. A PerkinElmer® Spectrum Two™ FTIR spectrometer
[0430] 55763132-1(PerkinElmer® Inc., USA) and the Dxcover (RTM) Platform Software (Dxcover Ltd., UK) automated spectral data acquisition were used to generate spectral data (16 co-added scans at 4 cm-1resolution, with 1 cm-1data spacing). Three spectra were collected for each sample well, resulting in nine replicates per patient.
[0431] The training dataset used for the development of the machine learning algorithm comprised 18,792 spectra, with 3,096 spectra labelled as poor quality (‘0’) and 15,696 labelled as good quality (‘T). This resulted in a prevalence of good quality spectra being about 84%. This data consisted of clinical serum data (n=17,442 spectra) and control data, with 900 LHPS spectra and 450 BSA spectra. LHPS were used as examples of positive controls, while BSA were used as examples of negative controls. All clinical serum spectra were labelled as either good or poor quality spectra after visual inspection. Other ratios of good to bad spectra may be employed.
[0432] Therefore, to train the machine learning algorithm, a classification is allocated to the reference spectra of the training data set using a binary score. In other embodiments, a graded (e.g. a percentage score) or other suitable classification may be allocated to the training data set. Whilst the classification was carried out by visual inspection, any suitable manner such as image processing could have been used instead. The classification is allocated prior to training the machine learning algorithm, but without wishing to be bound by theory, there may be other ways of doing this, or it may be unnecessary, depending on the properties of the testing method, the machine learning algorithm, and the hardware and software used.
[0433] The method of training the machine learning algorithm will comprise selecting the reference machine learning parameters to be used. These are selected based on knowledge of which spectral parameters influence the quality of the measured spectrum, and a list of potential candidates for the reference parameters can be found throughout this specification. An iterative process can be used to determine which reference machine learning parameters produce acceptable results, and those which may be superfluous or otherwise not required. The reference parameters will be selected based on the type of spectroscopy to be carried out, and without wishing to be bound by theory, it is thought by the present inventors to be possible to quickly retrain the machine learning algorithm by identifying new reference machine learning parameters for the new spectroscopic technique.
[0434] The method of training the machine learning algorithm comprises a plurality of pre-processing steps applied to some or all of the reference parameters of the reference spectrum or spectra, or to the reference spectra.
[0435] 55763132-1One applied pre-processing step is to cut the reference spectra to the 3500 cm-1to 1000 cm-1wavenumber region by removing upper and lower ends thereof. Another pre-processing step is to then apply a correction to the slope using a reference spectrum.
[0436] For the embodiment of Fig. 1, the pre-processing steps include selectively modifying a reference parameter of a reference spectrum by applying a modification function thereto. The modification function is a step function configured to modify the reference parameter for a first condition and to not modify the reference parameter for a second condition. This step is applied to all of the reference spectra, to modify only some of the reference machine learning parameters thereof. In the embodiment of Fig. 1, the following families of parameters were modified with the help of step-functions: “area_to_mean”, “dist_to_mean”, “sum_prom”, and “signal_to_noise”. Depending on the original value of the reference parameter, modification to the value may occur. For example, one particular reference parameter may return a value of between -1 and 1, but it is known that the effect on the quality of the spectrum of values between -1 and 0, of this parameter, is substantially the same, and therefore it may be desirable for the machine learning algorithm to treat parameter values between -1 and 0 as the same as a value of 0. This avoids anomalously low values from adversely affecting the quality test.
[0437] The step function is configured to: for obtained values of the reference parameter within the first condition, to attribute the same value for the modified reference parameter; and for obtained values within the second condition, to maintain the original value of the reference parameter. In other embodiments, the specific modification may be different than as described here.
[0438] The training of the algorithm of the embodiment of Fig. 1 therefore comprises applying four step functions to four reference parameters of the reference spectra.
[0439] In other embodiments, the pre-processing may comprise other functions, data processing, or the like, to modify the training data set of reference spectra prior to training the machine learning algorithm thereon.
[0440] The training of the algorithm comprises a first training step of the machine learning algorithm on a training portion of the training data set. The first training step comprises 51 iterations using a training portion selected such that, for at least two iterations, the training portion is different. The selection of the training portion is by substantially random selection from the training data set.
[0441] The first training step comprise testing the machine learning algorithm on a test portion of the training data set. The testing includes having the machine learning
[0442] 55763132-1algorithm make predictions on whether the reference spectra in the test portion are of sufficient quality. The step of testing is carried out for at least some of the iterations of the first training step. The test portion, for each iteration, excludes the training portion of that iteration.
[0443] Each training portion comprises 70% of the reference spectra of the training data set, and the test data makes up the remaining 30% of the training data set.
[0444] Each iteration of the first training step on a training portion includes a five-fold nested cross-validation (CV) process to train the algorithm on that training portion. The CV process includes, for each fold, splitting the training portion into a sub-training portion and sub-test portion, and training the algorithm on the sub-training portion and testing the algorithm on the sub-testing portion.
[0445] The first training step is an ensemble training method in which bagging or bootstrap aggregation is used, in which the selection of a reference spectrum for inclusion in the training portion is done with replacement, such that the same reference spectrum may be selected for inclusion more than once. Bagging is also used for the CV step of selecting each sub-training portion.
[0446] The machine learning algorithm comprises model hyper-parameters, which are set and adjusted during the first training step. The selection and adjustment of the model hyper-parameters contribute to the training of the machine learning algorithm.
[0447] The step of adjusting the hyper-parameters is carried out after an iteration of the first training step.
[0448] A second training step of training the machine learning algorithm is carried out after the first training step and after adjustment of the model hyper-parameters.
[0449] The second training step comprises training the machine learning algorithm on a training portion of the training data set using 51 iterations. In the second training step the training portion is a random selection from the training data set. In the second training step, the training portion is 70% of the reference spectra, and the testing portion is 30% of the training data set.
[0450] The second training step comprise testing the machine learning algorithm on the test portion of the training data set. The testing includes having the machine learning algorithm make predictions on whether the reference spectra in the test portion are of sufficient quality. The step of testing is carried out for all of the iterations of the second training step. The test portion, for each iteration, excludes the training portion of that iteration.
[0451] 55763132-1In the embodiment of Fig. 1 , the second training step is used to train and test the optimised algorithm from the first training step by keeping the model hyper-parameters fixed. It will be understood that in other embodiments, the first and / or second training steps may be different than as described here, and it is not essential to have multiple training steps. In other embodiments, it is envisaged that the data set used in the second training step may be different to the first training step.
[0452] Like the first training step, each iteration of the second training step on a training portion includes a five-fold nested CV to train the algorithm on that training portion. The CV includes, for each fold, splitting the training portion into a sub-training portion and sub-test portion, and training the algorithm on the sub-training portion and testing the algorithm on the sub-testing portion.
[0453] Like the first training step, the second training step is an ensemble training method comprising bagging or bootstrap aggregation, in which the selection of a reference spectrum for inclusion in the training portion is done with replacement, such that the same reference spectrum may be selected for inclusion more than once. The bagging is used for the CV step of selecting the, or each sub-training portion.
[0454] The method comprises a third training step of the machine learning algorithm carried out after the first training step and the second training step.
[0455] The third training step trained the machine learning algorithm on a training portion that is 100% of the training data set, comprising 51 iterations. Each iteration included a five-fold nested CV process, to train the algorithm. The CV included, for each fold, splitting the training portion (i.e. the entire training data set) into a sub-training portion and sub-test portion, and training the algorithm on the sub-training portion and testing the algorithm on the sub-testing portion.
[0456] The third training step is an ensemble training method comprising bagging or bootstrap aggregation used for the CV step of selecting the, or each sub-training portion.
[0457] It will be understood that in some embodiments, the third training step could be different than as described here.
[0458] The final model, a random forest model, and during the third training step, had the hyper-parameter mtry optimised which corresponds to the number of features randomly chosen in each split of the CV.
[0459] The minority class (‘fail’) was up-sampled during training.
[0460] The discriminatory threshold (tuned at 95% sensitivity) was chosen from the CV by spectrum Receiver Operating Characteristic (ROC) curves following an additional 51 re-iterations experiment.
[0461] 55763132-1The table below shows the mean performance results over the external test sets for the 51 re-iterations.
[0462]
[0463] The machine learning algorithm was further tested during full system verification using 22,410 spectra, of which 70 spectra were deemed poor quality and were failed by the algorithm. Figure 4 demonstrates a spectrum which has passed and a spectrum which has failed the quality test, obtained from the same sample, with the two spectra offset for ease of comparison.
[0464] The first training step, the second training step and / or the third training step each include a random forest decision tree method, which is an ensemble of decision trees and a supervised learning algorithm. It will be appreciated that this is one example of a machine learning algorithm and it is possible to adopt a different approach, either for training or the implementation of the invention.
[0465] The training method detailed here results in a machine learning algorithm that can determine the quality of negative and positive controls.
[0466] The processor element 2 is operable to adjust the threshold at which a spectrum is deemed of acceptable quality.
[0467] Although not shown here, the control testing method may comprise, by the processor element 2, highlighting and / or removing measured spectrum / spectra that are not of acceptable quality. The removal may be by deletion, archiving or the like.
[0468] The numbering of the training steps is intended to differentiate the training steps and not to convey any order or preference unless the context explicitly provides otherwise. It is possible in some examples that other steps may be applied in addition to those described here. Furthermore, it is in no way essential that all of the training steps are required, although some embodiments may make use of all the training steps described here. It is also envisaged that the specific order of the steps may be different than as described here.
[0469] The machine learning algorithm, for a mode of operation, is trained on reference spectrum / spectra corresponding to substantially the same type of sample. That is, the machine learning algorithm, for the second mode of operation, is trained on reference spectrum / spectra of a plurality of second samples, which are not identical but are
[0470] 55763132-1prepared in substantially the same way. The machine learning algorithm, for the third mode of operation, is likewise trained on reference spectrum / spectra of a plurality of third samples. It will be appreciated that for some embodiments, there may be one algorithm for the second mode, and another algorithm for the third mode.
[0471] One, some or all of the at least one reference machine learning parameter may be determined, at least in part, by the training of the machine learning algorithm. For example, training may help in deciding the specific values and ranges to consider adopting for the reference machine learning parameters.
[0472] For training, at least some of the reference spectra may be obtained from tests carried out using the same spectrometer 1 running the same mode of operation, or the reference spectra may be obtained, at least partially, from another source.
[0473] In the embodiments illustrated and described here, each spectrum parameter and each reference parameter, including the reference machine learning parameters, have a real value.
[0474] Some parts of the method of control testing are computer implemented, but there are embodiments of some aspects of the present disclosure that may be wholly computer implemented.
[0475] The steps of the method can be carried out in any order unless the context provides otherwise.
[0476] Figure 5 depicts a flow chart illustrating control testing of the spectrometer 1 followed by carrying out clinical diagnostic measurements using the spectrometer. First, the analysis instrument (the spectrometer 1 and associated components and software) is set up. Next, control testing is performed using control materials, and control data generated. If the quality tests produce acceptable results, the method proceeds to the clinical testing portion, and if not, a repeat of control testing is performed until acceptable.
[0477] Next, clinical diagnostic testing using the analysis instrument is carried out, to generate clinical data. The quality of the clinical data is checked, which may involve any suitable technique, and a diagnostic predication is generated.
[0478] Figure 5 therefore illustrates that the likely quality of clinical result is determined first by running control testing, before a clinical test is carried out, at least in some examples of a diagnostic method. It will be understood from a reading of the specification that steps may be omitted or added to the flow chart of Fig. 5, and that Fig. 5 shows an example method and it is not essential to carry out all of the steps in the manner shown therein and modifications may be readily made.
[0479] 55763132-1Although the disclosure has been described in terms of example embodiments as set forth above, it should be understood that these embodiments are illustrative only and that the claims are not limited to those embodiments. Those skilled in the art will be able to make modifications and alternatives in view of the disclosure, which are contemplated as falling within the scope of the appended claims. Each feature disclosed or illustrated in the present specification may be incorporated in any embodiments, whether alone or in any appropriate combination with any other feature disclosed or illustrated herein.
[0480] 55763132-1
Claims
CLAIMS:
1. A method of control testing for a non-targeted analysis spectroscopic measurement, the method comprising:(i) carrying out a first mode of control testing a spectrometer, independent of a subsequent non-targeted analysis diagnostic measurement to be carried out on a test sample by the spectrometer, the first mode comprising:providing a first, reference sample to the spectrometer, wherein the first, reference sample is configured to be independent of the test sample;obtaining, by a processor element, a first measured spectrum of the first, reference sample;determining if the first measured spectrum is of acceptable quality; and / or (ii) carrying out a second mode of negative control testing in respect of the subsequent diagnostic measurement to be carried out on the test sample by the spectrometer, the second mode comprising:providing a second, negative control sample to the spectrometer, wherein the second sample is configured to produce a different measured spectrum to a measured spectrum of the test sample;obtaining, by the processor element, a second measured spectrum of the second, negative control sample;determining if the second measured spectrum is of acceptable quality; and / or (iii) carrying out a third mode of positive control testing in respect of the subsequent diagnostic measurement, the third mode comprising:providing a third, positive control sample to the spectrometer, wherein the third sample is configured to produce a same or similar measured spectrum to the measured spectrum of the test sample;obtaining, by the processor element, a third measured spectrum of the third, positive control sample;determining if the third measured spectrum is of acceptable quality.
2. The method of claim 1 , wherein the method comprises carrying out the first mode, the second mode and the third mode.55763132-13. The method of claim 1 or claim 2, wherein the, or each step of determining the quality of the first measured spectrum, the second measured spectrum and / or the third measured spectrum comprises:identifying, by the processor element, at least one spectrum parameter of the measured spectrum; andcomparing, by the processor element, the at least one spectrum parameter against at least one reference parameter.
4. The method of any preceding claim, wherein the processor element is connected to or is part of the spectrometer.
5. The method of any preceding claim, wherein the step of determining the quality of the first measured spectrum, the second measured spectrum, and / or the third measured spectrum is carried out automatically by the processor element.
6. The method of any preceding claim, wherein the first, reference sample is a non-biological control sample.
7. The method of any preceding claim, wherein the second, negative control sample comprises human or animal biological material.
8. The method of claim 7, wherein the second, negative control sample comprises one or more components of and / or one or more comparable components to the third sample and / or to the test sample.
9. The method of claim 7 or claim 8, wherein the second, negative control sample includes blood, plasma, serum, bovine serum albumin (BSA), a BSA solution, bovine serum plasma, and / or a bovine serum plasma solution.
10. The method of any preceding claim, wherein the third, positive control sample comprises human or animal biological material.
11. The method of claim 10, wherein the third, positive control sample comprises a carrier component, such as blood, plasma or serum, and the third, positive control sample comprises substantially the same carrier component as the test sample.55763132-112. The method of claim 10 or claim 11, wherein the third, positive control sample includes blood, serum, plasma, lyophilised human pooled serum (LHPS), a LHPS solution, lyophilised human pooled plasma, and / or a lyophilised human pooled plasma solution.
13. The method of any of claims 3 to 12, wherein at least one, or some, or all of the reference parameter(s) may be different between at least two of the first mode, the second mode and the third mode.
14. The method of any of claims 3 to 13, wherein, in at least one mode of control testing, at least one of the steps of: comparing the at least one spectrum parameter against the at least one reference parameter; and determining the quality of the measured spectrum comprises using a machine learning algorithm and at least one reference machine learning parameter as the reference parameter.
15. The method of claim 14, wherein the first mode uses at least one reference parameter and a numerical comparison thereto; and wherein the second mode and / or the third mode uses the machine learning algorithm and the at least one reference machine learning parameter.
16. The method of any preceding claim, wherein the processor element is operable to adjust the threshold at which a spectrum is deemed of acceptable quality.
17. The method of any of claims 3 to 16, wherein the at least one reference parameter includes one or more of the following:the intensity values for a range of frequencies;an intensity value for a frequency;a location of one or more expected peak values or trough values;a peak intensity value or trough intensity value (maximum value or minimum value) at one or more frequencies;a prominence of one or more peak intensity values or trough intensity values for a range of frequencies;a signal to noise ratio of an intensity value, or intensity values of the spectrum, which may include the signal to noise ratio of a peak intensity value or trough intensity55763132-1value against another part of the spectrum, such as a non-signal part or non-peak / trough part of the spectrum; and / ora relative peak to trough value within a range of frequencies.
18. The method of any of claims 15 to 17, wherein the numerical comparison comprises at least one of: comparing a measured parameter to a reference parameter using an arithmetic operation, such as subtraction, or a vector operation such as obtaining an Euclidean norm of the difference of a measured spectrum parameter to a reference parameter.
19. The method of any of claims 14 to 18, wherein the machine learning algorithm is a decision tree method, which may be an ensemble of decision trees.
20. The method of claim 19, wherein the machine learning algorithm is a supervised learning algorithm.
21. The method of claim 19 or claim 20, wherein the decision tree method is a random forest model.
22. The method of any of claims 14 to 21 , wherein the at least one reference machine learning parameter includes one or more of:a location of a peak intensity value or trough intensity value within a range of frequencies;a location of a maximum or minimum intensity value within a range of frequencies;a value of a peak or trough for a range of frequencies;a location of a peak intensity value or trough intensity value within a region of the spectrum associated with a substance (e.g. within an absorbance band of the substance);a location of a local or global peak or global trough within a range of frequencies; the intensity value of the local or global peak / trough;spectral intensity values for a range of frequencies;a transformation parameter used to apply a transformation or correction to the reference spectrum or reference spectra, or to the reference machine learning parameter(s);55763132-1a ratio of the intensity value of a local maximum within a range of frequencies to a local minimum within a range of frequencies, wherein the ranges may be different or substantially the same;a signal to noise ratio, which may include the signal to noise ratio of a peak intensity value or trough intensity value against another part of the spectrum, such as a non-signal part or non-peak / trough part of the spectrum;a prominence of one or more peak values or trough values;a sum of prominences of a local minima and maxima within a range of frequencies;a width of a peak or trough for a range of frequencies, which may be for a substance band.
23. The method of any of claims 14 to 22, wherein the step of comparing the at least one spectrum parameter to the at least one reference machine learning parameter comprises measuring one or more of:a peak or trough x-shift (frequency shift) of the measured spectrum parameter against the reference machine learning parameter;a peak or trough y-shift (difference in intensity) of the measured spectrum parameter against the reference machine learning parameter;a numerical comparison or arithmetic operation between the measured spectrum parameter and the reference machine learning parameter;a calculation of the degree of similarity to expected value(s), which may be a percentage value or the like;comparison to mean values, which may be a mean reference parameter obtained from a plurality of reference parameters, where the mean may be for a range of frequencies;an area difference to the mean, which may be for a range of frequencies;a cosine distance to the mean, which may be for a range of frequencies;a 1-norm distance to the mean, which may be for a range of frequencies; the number of intensity values that are outside of n standard deviations from the mean, where n can be e.g. 1, 2, 3, or any suitable value;a ratiometric difference between a local maximum and a local minimum with a range of frequencies;a sum of two or more prominences of local minimum / minima and local maximum / maxima in a range of frequencies;55763132-1determination if a parameter is above or below a threshold value, such as assigning a “1” if the width of an intensity peak is below a threshold width, and a “0” if above the threshold width.
24. An apparatus for non-targeted analysis diagnostic measurement of a test sample, the apparatus comprising:a spectrometer operable to receive a sample for measurement thereof;a processor element configured, in use, to obtain a measured spectrum of the sample;wherein the apparatus is operable to carry out:(i) a first mode of control testing the spectrometer on a first, reference sample; and / or(ii) a second mode of negative control testing on a second, negative control sample; and / or(iii) a third mode of positive control testing on a third, positive control sample.
25. A diagnostic method comprising:running a method of control testing according to any of claims 1 to 23; andrunning a diagnostic test using the spectrometer.55763132-1