A non-invasive diagnosis system

A non-invasive breath analysis system using machine learning on VOCs effectively classifies diabetes status, addressing the limitations of current invasive and inaccurate tests, enabling early and accurate T2DM detection.

WO2026154262A1PCT designated stage Publication Date: 2026-07-23NANO-NOSE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NANO-NOSE LTD
Filing Date
2026-01-15
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Current methods for diagnosing Type 2 diabetes mellitus (T2DM) are invasive, costly, and lack specificity, with existing non-invasive tests producing high false positives and negatives, and there is a need for an affordable, non-invasive, and accurate method for early detection.

Method used

A non-invasive diagnosis system using a device to collect breath samples, analyze VOCs with a spectrometer, and apply machine learning models to classify individuals as healthy, prediabetic, or with Type 1 or Type 2 diabetes based on combinations of VOCs, utilizing a training dataset that includes glucose levels, HbA1c tests, and OGTT results.

Benefits of technology

The system provides accurate and early detection of T2DM with high sensitivity and specificity, allowing for self-monitoring by patients, reducing the risk of complications and improving diagnostic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2026050040_23072026_PF_FP_ABST
    Figure GB2026050040_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A non-invasive diagnosis system, which comprises: a device for collecting a breath sample from a user, which comprises a mouthpiece; a spectrometer for measuring Volatile Organic Compounds, VOCs, in the breath sample; and a processor, wherein the processor is configured to analyse the measured VOCs using one or more machine learning models for classifying the user, based on one or more combinations of the measured VOCs, as healthy, prediabetic, Type 1 or Type 2 diabetes mellitus.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] P42655GB

[0002] A non-invasive diagnosis system

[0003] The present disclosure relates to a non-invasive diagnosis system, in particular to such a system for determining whether a user is healthy, prediabetic, Type 1 or Type 2 diabetes mellitus (T1DM, T2DM).

[0004] In 2013 there were an estimated 382 million people worldwide living with diabetes in one form or another, and by 2019 this had increased to 463 million. It is anticipated that by the year 2030, 578 million people will have diabetes which will increase to 700 million in 2045. The largest increase in diabetes cases over the next 20 years is expected to come from countries transitioning from low to middle income.

[0005] In the United Kingdom, incidence rates have more than doubled from 1.4 million to 3.8 million since 1996, and an estimated 4.7 million people will live with diabetes, of which 1 million are undiagnosed, by 2025.

[0006] There are clear benefits to early diagnosis, in particular, reducing the number of people that will experience diabetes-related complications such as sight loss, amputation, kidney failure, stroke and heart disease.

[0007] Glucose intolerance is caused by insulin resistance and / or defective insulin secretion. The glucose tolerance scale is a continuous spectrum from normal to T2DM, wherein glycaemic limits have been introduced to highlight people that are at an increased risk of developing T2DM. Impaired glucose regulation ('prediabetes’) covers blood glucose levels that are above the normal range but are not high enough for the diagnosis of T2DM. Impaired glucose regulation is asymptomatic and often remains undiagnosed until a transition to T2DM occurs and symptoms begin.

[0008] Early identification of glucose intolerance is difficult. A standardised 75g oral glucose tolerance test (OGTT) has been used for decades for the diagnosis of varying degrees of glucose intolerance. However, the cut-offs used for diagnosis have changed over time. More recently, the use of a test that measures a bloodsample for levels of glycated haemoglobin (HbA1c) has been advocated for the diagnosis of diabetes. Measurement of HbA1c offers some advantages over conducting OGTTs, such as not requiring an overnight fast and being cheaper than performing an OGTT. However, there are drawbacks to using HbA1c to diagnose T2DM. The results can be affected by haemolysis and other conditions that affect red cell turnover rates. Several online risk score calculators have now been developed to identify individuals at higher risk of developing T2DM without the need for conventional invasive biochemical tests. These risk-assessment tools have low specificity resulting in large numbers of false positives resulting in additional testing being carried out unnecessarily. People with HbA1c between 42-48 mmol / mol in the UK are considered high risk (prediabetes) for developing T2DM.

[0009] By the time HbA1c has risen above 'normal 'levels, i.e. >42 mmol / mol, they may have had some level of glucose intolerance for a considerable time.

[0010] There is currently no simple, affordable, and non-invasive test for early T2DM.

[0011] More than 3000 VOCs have been detected in the breath of humans based on earlier third-party studies. Prior research has considered whether certain detected combinations of these VOCs may allow for differentiating between individuals with diabetes and healthy controls in specific contexts for the diagnosis of diabetes or prediabetes.

[0012] Possible useful combinations were identified as follows:

[0013] • 2,3,4-trimethylhexane, 2,6,8-trimethyldecane, isopropanol, tridecane and undecane;

[0014] • Acetone, butanol, dimethyl sulphide, isoprene, pyridine and compounds identified by Proton-transfer-reaction mass spectrometry (PTR-MS) as masses 36, 49 and 95; and

[0015] • Acetone, butanol, isoprene and compounds identified by PTR-MS as masses 49, 126, 135, 140 and 157.Traditional models (e.g., regression type) were implemented, using a ranking method to assign a level of importance to these VOCs in an effort to make a determination between healthy, pre- and type 2 diabetics.

[0016] However, results were not consistent. No effective method of diagnosis or classification was determined.

[0017] The present invention arose as a result of work aimed at developing an earlier, quicker, easier way to diagnose T2DM.

[0018] Representative features are set out in the following clauses, which stand alone or may be combined, in any combination, with one or more features disclosed in the text and / or drawings of the specification.

[0019] According to the present invention, in a first aspect, there is provided a non-invasive diagnosis system, which comprises: a device for collecting a breath sample from a user, which comprises a mouthpiece; a spectrometer for measuring Volatile Organic Compounds, VOCs, in the breath sample; and a processor, wherein the processor is configured to analyse the measured VOCs using one or more machine learning models for classifying the user, based on one or more combinations of the measured VOCs, as healthy, prediabetic, Type 1 or Type 2 diabetes mellitus.

[0020] Preferably, the one or more machine learning models comprise a model generated using a labelled training data set, which comprises a plurality of data points, wherein each data point corresponds to an individual and is associated with one of the labels healthy, prediabetic, Type 1 or Type 2 diabetes mellitus for the individual, and each data point comprises the measured VOCs from one or more breath samples from the individual.

[0021] According to the present invention in a further aspect, there is provided a method for classifying individuals using machine learning, the method comprising: training a machine learning model using a labelled training dataset, which comprises a plurality of data points, wherein each data point corresponds to an individual and is associated with one of the labels healthy, prediabetic, Type 1 or Type 2 diabetesmellitus for the individual, and each data point comprises the measured VOCs from one or more breath samples from the individual; applying the trained machine learning model to classify an individual based on one or more breath samples taken from the individual; and providing the classification result.

[0022] Preferably, each data point is further associated with a glucose level of the individual.

[0023] Each data point may comprise the result of an HbA1c test performed on the individual.

[0024] Each data point may comprise the results of an oral glucose tolerance test, OGTT, performed on the individual.

[0025] Preferably, each data point further comprises additional data on the individual, including one or more or all of age, sex, weight, height and body mass index.

[0026] The one or more machine learning models may comprise a model generated using training data that comprises a plurality or all of the following VOCs: Norflurane, Carbonyl sulfide, Dimethyl ether, Acetaldehyde, Butane, 1 -isocyano-, Ethanol, Isobutane, Acetonitrile, Methyl isocyanide, Acetone, Isopropyl Alcohol, Furan, Isoprene, Methylene chloride, Carbon disulfide, Acetic acid, n-Hexane, Furan, 2-methyl-, Tetrahydrofuran, Butanal, 3-methyl-, Benzene, 1 -Butanol, Butanal, 2-methyl-, Thiophene, 2-Propanol, 1 -methoxy-, 2-Pentanone, Heptane, Pentanal, Furan, 2,5-dimethyl-, 1,4-Dioxane, 1 -Pentene, 2,4,4-trimethyl-, Butanenitrile, 3-methyl-, Methyl Isobutyl Ketone, Toluene, Thiophene, 3-methyl-, 1 -Octene, 2-Hexanone, Octane, Hexanal, Maleic anhydride, 1H-Pyrrole, 2,5-dimethyl-, 1-Hexanol, Lactic acid, Benzene, 1 ,3-dimethyl-, Pentanoic acid, 1 -Nonene, 2-Heptanone, 4-Cyclopentene-1 ,3-dione, Styrene, Heptanal, Butyrolactone, Disulfide, methyl 2-propenyl, 2-Heptanone, 4-methyl-, 2,5-Furandione, dihydro-3-methylene-2(3H)-Furanone, dihydro-5-methyl-, Hexanoic acid, 1 -Nonanol, Benzaldehyde, Phenol, 5-Hepten-2-one, 6-methyl-, 2-Heptanone, 4,6-dimethyl-, Furan, 2-pentyl-, Benzonitrile, Decane, Hexanoic acid, ethyl ester, Octanal, Benzofuran, 1 -Hexanol, 2-ethyl-, Cyclohexene, 1-methyl-4-(1 -methylethenyl)-, (S)-, Benzeneacetaldehyde,Benzaldehyde, 2-hydroxy-, 2H-Pyran-2-one, tetrahydro-, 1 -Octanol, Acetophenone, Benzaldehyde, 2-methyl-, 4-Undecene, (E)-, Diallyl disulphide, Undecane, Nonanal, Benzofuran, 2-methyl-, 1 -Propanol, 3,3'-oxybis-, Benzoic acid, Octanoic acid, Alkanoll, 1 ,2-Propanedione, 1 -phenyl-, 1 -Dodecene, 2-Decanone, Dodecane, Ethanone, 1-(3-methylphenyl)-, Naphthalene, Decanal, 2-Propanol, 1-(2-butoxy-1-methylethoxy)-, Benzothiazole, Nonanoic acid, 1 -Decanol, Tridecane, Undecanal, n-Decanoic acid, Pentanoic acid, 2,2,4-trimethyl-3-hydroxy-, isobutyl ester, 1(3H)-Isobenzofuranone, Tetradecane, Dodecanal, Ethanone, 2, 2-dihydroxy-1 -phenyl-, 1 ,3-lsobenzofurandione, 4-methyl-, 5,9-Undecadien-2-one, 6,10-dimethyl-, (E)-, 4-Methylphthalic anhydride, Dodecane, 1 -chloro-, Alkenel, 1 -Pentadecene, o-Cyanobenzoic acid, Pentadecane, 2,4-Di-tert-butylphenol, Tridecanal, 1-Penten-3-one, 1-(2,6,6-trimethyl-2-cyclohexen-1-yl)-, Stilbene, o-Hydroxybiphenyl, Lilial, Dodecanoic acid, Phenylmaleic anhydride, Pentadecane, 3-methyl-, Pentanoic acid, 2,2,4-trimethyl-3-carboxyisopropyl, isobutyl ester, Hexadecane, Phenol, 2-cyclohexyl-, Tetradecanal, 4-Cyclopentene-1 ,3-dione, 4-phenyl-, Sesquiterpenel , Benzophenone, Cyclopentaneacetic acid, 3-oxo-2-pentyl-, methyl ester, Heptadecane, o-Anisic acid, 4-nitrophenyl ester, Tetradecanoic acid, Alkyl phenol, Octanal, 2-(phenylmethylene)-, Ethanone, 1 ,2-diphenyl-, 9H-Fluoren-9-one, Octadecane, Hexadecane, 2,6,10,14-tetramethyl-, Benzenesulfonamide, N-butyl-, 2-Ethylhexyl salicylate, Isopropyl myristate, Benzene, 1 ,1'-[1 ,2-ethanediylbis(oxy)]bis-, Pentadecanoic acid, 10-Methylanthracene-9-carboxaldehyde, Oxacyclotetradecane-2, 11 -dione, 13-methyl-, 7-Acetyl-6-ethyl-1,1,4,4-tetramethyltetralin, Cetene, Naphthalic anhydride, Nonadecane, Benzoic acid, 2-hydroxy-, phenylmethyl ester, (E, E)-7, 11 ,15-Trimethyl-3-methylene-hexadeca-1 ,6,10,14-tetraene, Hexadecanoic acid, methyl ester, 9-Hexadecenoic acid, Nonadecane, 4-methyl-, n-Hexadecanoic acid, Diphenyl sulfone, Dibutyl phthalate, Eicosane, Isopropyl palmitate, Squalene, Hexanoic acid, 2-ethyl-, hexadecyl ester, Eicosane, 3-methyl-, 1 -Octadecene, Methanone, 2-benzofuranylphenyl- Heneicosane, 2,3-Diphenylmaleic anhydride, Oleic Acid, Octadecanoic acid, Docosane, Benzoic acid, tetradecyl ester, p-Terphenyl, Tricosane, 2-methyl-, [1 ,1':3',1"-Terphenyl]-2'-ol, Tricosane, Hexacosane, Ethane, Pentane, 2-Methylpentane, 3-Methylpentane, Formaldehyde, Propanal, Malondialdehyde, 4-Hydroxynonenal, 2-Butanone (methyl ethyl ketone), 3-Pentanone, Cyclohexanone, Acetoin, Diacetyl (2,3-butanedione), Methanol, 1-Propanol, 1 -Pentanol, Formic acid, Propionic acid, Butyric acid, Isovaleric acid,Methyl acetate, Ethyl acetate, Hydrogen sulfide, Methanethiol, Dimethyl sulfide, Dimethyl trisulfide, Ammonia, Nitric oxide, Nitrogen dioxide, Trimethylamine, Pyridine, Indole, Skatole, p-Cresol, o-Cresol, lndole-3-acetaldehyde, Limonene, Pinene, Geranyl acetone.

[0027] Each data point may comprise a plurality or all of these VOCs.

[0028] The training data set may exclude saturated hydrocarbons.

[0029] The one or more machine learning models may comprise a model generated using training data that comprises a plurality or all of the following VOCs: Norflurane, Carbonyl sulfide, Dimethyl ether, Acetaldehyde, Butane, 1 -isocyano-, Ethanol, Acetonitrile, Methyl isocyanide, Acetone, Isopropyl Alcohol, Furan, Isoprene, Methylene chloride, Carbon disulfide, Acetic acid, Furan, 2-methyl-, Tetrahydrofuran, Butanal, 3-methyl-, Benzene, 1 -Butanol, Butanal, 2-methyl-, Thiophene, 2-Propanol, 1 -methoxy-, 2-Pentanone, Pentanal, Furan, 2,5-dimethyl-, 1,4-Dioxane, 1 -Pentene, 2,4,4-trimethyl-, Butanenitrile, 3-methyl-, Methyl Isobutyl Ketone, Toluene, Thiophene, 3-methyl-, 1 -Octene, 2-Hexanone, Hexanal, Maleic anhydride, 1H-Pyrrole, 2,5-dimethyl-, 1 -Hexanol, Lactic acid, Benzene, 1 ,3-dimethyl-, Pentanoic acid, 1 -Nonene, 2-Heptanone, 4-Cyclopentene-1 ,3-dione, Styrene, Heptanal, Butyrolactone, Disulfide, methyl 2-propenyl, 2-Heptanone, 4-methyl-, 2,5-Furandione, dihydro-3-methylene-, 2(3H)-Furanone, dihydro-5-methyl-, Hexanoic acid, 1 -Nonanol, Benzaldehyde, Phenol, 5-Hepten-2-one, 6-methyl-, 2-Heptanone, 4,6-dimethyl-, Furan, 2-pentyl-, Benzonitrile, Hexanoic acid, ethyl ester, Octanal, Benzofuran, 1 -Hexanol, 2 -ethyl-, Cyclohexene, 1-methyl-4-(1 -methylethenyl)-, (S)-, Benzeneacetaldehyde, Benzaldehyde, 2-hydroxy-, 2H-Pyran-2-one, tetrahydro-, 1-Octanol, Acetophenone, Benzaldehyde, 2-methyl-, 4-Undecene, (E)-, Diallyl disulphide, Nonanal, Benzofuran, 2-methyl-, 1 -Propanol, 3,3'-oxybis-, Benzoic acid, Octanoic acid, Alkanoll, 1,2-Propanedione, 1 -phenyl-, 1 -Dodecene, 2-Decanone, Ethanone, 1-(3-methylphenyl)-, Naphthalene, Decanal, 2-Propanol, 1-(2-butoxy-1-methylethoxy)-, Benzothiazole, Nonanoic acid, 1 -Decanol, Undecanal, n-Decanoic acid, Pentanoic acid, 2,2,4-trimethyl-3-hydroxy-, isobutyl ester, 1(3H)-Isobenzofuranone, Dodecanal, Ethanone, 2, 2-dihydroxy-1 -phenyl-, 1,3-Isobenzofurandione, 4-methyl-, 5,9-Undecadien-2-one, 6,10-dimethyl-, (E)-, 4-Methylphthalic anhydride, Dodecane, 1 -chloro-, Alkenel, 1 -Pentadecene, o-Cyanobenzoic acid, 2,4-Di-tert-butylphenol, Tridecanal, 1-Penten-3-one, 1 -(2,6,6-trimethyl-2-cyclohexen-1-yl)-, Stilbene, o-Hydroxybiphenyl, Lilial, Dodecanoic acid, Phenylmaleic anhydride, Pentanoic acid, 2,2,4-trimethyl-3-carboxyisopropyl, isobutyl ester, Phenol, 2-cyclohexyl-, Tetradecanal, 4-Cyclopentene-1 ,3-dione, 4-phenyl-, Sesquiterpenel , Benzophenone, Cyclopentaneacetic acid, 3-oxo-2-pentyl-, methyl ester, o-Anisic acid, 4-nitrophenyl ester, Tetradecanoic acid, Alkyl phenol, Octanal, 2-(phenylmethylene)-, Ethanone, 1 ,2-diphenyl-, 9H-Fluoren-9-one, Benzenesulfonamide, N-butyl-, 2-Ethylhexyl salicylate, Isopropyl myristate, Benzene, 1,1'-[1,2-ethanediylbis(oxy)]bis-, Pentadecanoic acid, 10-Methylanthracene-9-carboxaldehyde, Oxacyclotetradecane-2,11 -dione, 13-methyl-, 7-Acetyl-6-ethyl-1 ,1 ,4,4-tetramethyltetralin, Cetene, Naphthalic anhydride, Benzoic acid, 2-hydroxy-, phenylmethyl ester, (E, E)-7, 11 ,15-Trimethyl-3-methylene-hexadeca-1 ,6,10,14-tetraene, Hexadecanoic acid, methyl ester, 9-Hexadecenoic acid, n-Hexadecanoic acid, Diphenyl sulfone, Isopropyl palmitate, Squalene, Hexanoic acid, 2-ethyl-, hexadecyl ester, 1 -Octadecene, Methanone, 2-benzofuranylphenyl-, 2,3-Diphenylmaleic anhydride, Oleic Acid, Octadecanoic acid, Benzoic acid, tetradecyl ester, p-Terphenyl, [1 ,1':3',1"-Terphenyl]-2'-ol, Formaldehyde, Propanal, Malondialdehyde, 4-Hydroxynonenal, 2-Butanone (methyl ethyl ketone), 3-Pentanone, Cyclohexanone, Acetoin, Diacetyl (2,3-butanedione), Methanol, 1-Propanol, 1 -Pentanol, Formic acid, Propionic acid, Butyric acid, Isovaleric acid, Methyl acetate, Ethyl acetate, Hydrogen sulfide, Methanethiol, Dimethyl sulfide, Dimethyl trisulfide, Ammonia, Nitric oxide, Nitrogen dioxide, Trimethylamine, Pyridine, Indole, Skatole, p-Cresol, o-Cresol, lndole-3-acetaldehyde, Limonene, Pinene, Geranyl acetone.

[0030] Each data point may comprise a plurality or all of these VOCs.

[0031] The training data set may exclude all hydrocarbons.

[0032] The one or more machine learning models may comprise a model generated using training data that comprises a plurality or all of the following VOCs: Norflurane, Carbonyl sulfide, Dimethyl ether, Acetaldehyde, Butane, 1 -isocyano-, Ethanol, Acetonitrile, Methyl isocyanide, Acetone, Isopropyl Alcohol, Furan, Methylenechloride, Carbon disulfide, Acetic acid, Furan, 2-methyl-, Tetrahydrofuran, Butanal, 3-methyl-, Benzene, 1 -Butanol, Butanal, 2-methyl-, Thiophene, 2-Propanol, 1-methoxy-, 2-Pentanone, Pentanal, Furan, 2,5-dimethyl-, 1,4-Dioxane, Butanenitrile, 3-methyl-, Methyl Isobutyl Ketone, Toluene, Thiophene, 3-methyl-, 2-Hexanone, Hexanal, Maleic anhydride, 1H-Pyrrole, 2,5-dimethyl-, 1 -Hexanol, Lactic acid, Benzene, 1 ,3-dimethyl-, Pentanoic acid, 2-Heptanone, 4-Cyclopentene-1 ,3-dione, Styrene, Heptanal, Butyrolactone, Disulfide, methyl 2-propenyl, 2-Heptanone, 4-methyl-, 2,5-Furandione, dihydro-3-methylene-, 2(3H)-Furanone, dihydro-5-methyl-, Hexanoic acid, 1 -Nonanol, Benzaldehyde, Phenol, 5-Hepten-2-one, 6-methyl-, 2-Heptanone, 4,6-dimethyl-, Furan, 2-pentyl-, Benzonitrile, Hexanoic acid, ethyl ester, Octanal, Benzofuran, 1 -Hexanol, 2-ethyl-, Benzeneacetaldehyde, Benzaldehyde, 2-hydroxy-, 2H-Pyran-2-one, tetrahydro-, 1 -Octanol, Acetophenone, Benzaldehyde, 2-methyl-, Diallyl disulphide, Nonanal, Benzofuran, 2-methyl-, 1 -Propanol, 3,3'-oxybis-, Benzoic acid, Octanoic acid, Alkanoll, 1,2-Propanedione, 1 -phenyl-, 2-Decanone, Ethanone, 1-(3-methylphenyl)-, Naphthalene, Decanal, 2-Propanol, 1-(2-butoxy-1-methylethoxy)-, Benzothiazole, Nonanoic acid, 1 -Decanol, Undecanal, n-Decanoic acid, Pentanoic acid, 2,2,4-trimethyl-3-hydroxy-, isobutyl ester, 1(3H)-Isobenzofuranone, Dodecanal, Ethanone, 2, 2-dihydroxy-1 -phenyl-, 1,3-Isobenzofurandione, 4-methyl-, 5,9-Undecadien-2-one, 6,10-dimethyl-, (E)-, 4-Methylphthalic anhydride, Dodecane, 1 -chloro-, o-Cyanobenzoic acid, 2,4-Di-tert-butylphenol, Tridecanal, 1 -Penten-3-one, 1-(2,6,6-trimethyl-2-cyclohexen-1-yl)-, Stilbene, o-Hydroxybiphenyl, Lilial, Dodecanoic acid, Phenylmaleic anhydride, Pentanoic acid, 2,2,4-trimethyl-3-carboxyisopropyl, isobutyl ester, Phenol, 2-cyclohexyl-, Tetradecanal, 4-Cyclopentene-1 ,3-dione, 4-phenyl-, Benzophenone, Cyclopentaneacetic acid, 3-oxo-2-pentyl-, methyl ester, o-Anisic acid, 4-nitrophenyl ester, Tetradecanoic acid, Alkyl phenol, Octanal, 2-(phenylmethylene)-, Ethanone, 1 ,2-diphenyl-, 9H-Fluoren-9-one, Benzenesulfonamide, N-butyl-, 2-Ethylhexyl salicylate, Isopropyl myristate, Benzene, 1,1'-[1,2-ethanediylbis(oxy)]bis-Pentadecanoic acid, 10-Methylanthracene-9-carboxaldehyde, Oxacyclotetradecane-2, 11 -dione, 13-methyl-, 7-Acetyl-6-ethyl-1,1,4,4-tetramethyltetralin, Cetene, Naphthalic anhydride, Benzoic acid, 2-hydroxy-, phenylmethyl ester, Hexadecanoic acid, methyl ester, 9-Hexadecenoic acid, n-Hexadecanoic acid, Diphenyl sulfone, Isopropyl palmitate, Hexanoic acid, 2-ethyl-, hexadecyl ester, Methanone, 2-benzofuranylphenyl-, 2,3-Diphenylmaleic anhydride, Oleic Acid, Octadecanoic acid,Benzoic acid, tetradecyl ester, p-Terphenyl, [1 ,1':3',1"-Terphenyl]-2'-ol, Formaldehyde, Propanal, Malondialdehyde, 4-Hydroxynonenal, 2-Butanone (methyl ethyl ketone), 3-Pentanone, Cyclohexanone, Acetoin, Diacetyl (2,3-butanedione), Methanol, 1 -Propanol, 1 -Pentanol, Formic acid, Propionic acid, Butyric acid, Isovaleric acid, Methyl acetate, Ethyl acetate, Hydrogen sulfide, Methanethiol, Dimethyl sulfide, Dimethyl trisulfide, Ammonia, Nitric oxide, Nitrogen dioxide, Trimethylamine, Pyridine, Indole, Skatole, p-Cresol, o-Cresol, lndole-3-acetaldehyde, Geranyl acetone.

[0033] Each data point may comprise a plurality or all of these VOCs.

[0034] The training data set may include at least 10 VOCs, preferably at least 20 VOCs, and more preferably at least 30 VOCs.

[0035] The one or more machine learning models may comprise one or more machine learning models selected from: K-nearest neighbour, K-means, Gaussian Mixture Model, Support Vector Machine, Convolutional Neural Network, Random Forest Classifier, multi-layer perceptron, linear model, Extra Trees Regressor.

[0036] Alternative discriminative and generative models used for regression / classification may also be used.

[0037] The training dataset may be updated or augmented with real or synthetic data. The machine learning model may be improved or retrained based on user-provided corrections or newly labelled data.

[0038] The term synthetic data is intended to cover artificially generated data that mimics the characteristics and structure of real-world data, as will be appreciated by those skilled in the art.

[0039] Non-limiting embodiments of the invention will now be discussed with reference to the following drawings:

[0040] Figure 1 shows a schematic illustration of a system according to an embodiment of the present invention; andFigure 2 shows a flow chart summarising methods and processes of analyses according to one or more embodiments of the present invention.

[0041] The present invention relates to a non-invasive diagnosis system, as seen schematically in Figure 1 , which comprises: a device 1 for collecting a breath sample from a user 2, which comprises a mouthpiece 3; a spectrometer 4 for measuring Volatile Organic Compounds, VOCs, in the breath sample; and a processor 5, wherein the processor 5 is configured to analyse the measured VOCs using one or more machine learning models for classifying the user, based on one or more combinations of the measured VOCs, as healthy, prediabetic, Type 1 or Type 2 diabetes mellitus.

[0042] Testing exhaled breath for biomarkers of diabetes is a non-invasive method of diagnosing and monitoring the disease, and the simplicity of the method as detailed herein offers clear scope for use by patients unsupervised, bringing numerous advantages and allowing patients to easily monitor their condition.

[0043] Biomarker technologies detect gas phase VOCs in the exhaled breath of the patient. Identification of the relevant compounds and concentration levels are then used to monitor the ketotic state of diabetic and fasting individuals, and may further be used to measure glucose levels.

[0044] Among the challenges solved by the present inventors has been the identification of the VOCs (in exhaled breath) that can indicate the presence of diabetes with sufficiently high sensitivity and specificity, and which can be detected by technologies that may be miniaturised and produced at affordable cost.

[0045] As noted, more than 3000 VOCs have been detected in the breath of humans. The biomarkers fall into two main categories, endogenous, which occur naturally in the breath, and exogenous, which are exhaled after ingestion of a labelled tracer molecule. The focus of the present disclosure is on endogenous VOCs, as these are more suitable to a point-of-care diagnostic test, or for regular monitoring of people known to have diabetes.Without limitation, in development of the present invention, two mechanisms have been identified for the extraction and detection of the presence of different biomarkers.

[0046] Firstly, a methodology suited for use, for example, in laboratories, which operates under vacuum has been used for analysis of a breath sample using gas chromatography with mass spectrometry (GC-MS) operating below atmospheric pressure using nitrogen.

[0047] Secondly, a methodology suited for use, for example, outside of the laboratory, using ion mobility spectrometry (IMS), or similar analysis technologies, and operating at atmospheric pressure.

[0048] It should be appreciated that the form of spectrometer for analysis of the breath sample is not, however, to be limited to the above mentioned examples. Any suitable spectrometer capable of analysing the desired VOCs may be implemented, as will be readily appreciated by those skilled in the art.

[0049] The methodology for analysing the sample in the laboratory has good sensitivity and specificity, has been determined to be effective for identifying individuals who are healthy, pre-diabetic or with type 2 diabetes mellitus (T2DM), and was used, initially, to identify certain combinations of compounds.

[0050] However, certain combinations of these biomarkers are hard to use in the field, according to the second methodology, under the operating conditions. For example, isoprene is a hydrocarbon (and perceived previously to be an important biomarker), which is hard to ionise at atmospheric pressure.

[0051] Moreover, because human breath is ‘thick’, diverse and specific molecules in the air will be affected by surrounding elements.

[0052] Accordingly, the approach adopted by the present inventors has been to move away from the use of a ranking of a few markers that if present, or present in a given order, will lead to a successful diagnosis, and rather towards the considerationof multiple combinations of multiple markers for different scenarios that can lead to a successful diagnosis. This represents a difference between a ranking method and Al modelling of a multidimensional space in which, as discussed, one or more machine learning models are used for classifying the user, based on one or more combinations of the measured VOCs.

[0053] Methods according to the invention, as detailed further below, compare multiple combinations with multiple combinations and can achieve very high accuracy in a range of scenarios / using a range of models, as exemplarily presented in Figure 2, including, utilising data sets with all VOCs included, using data sets with saturated hydrocarbons removed, and using data sets with all hydrocarbons removed. This is a key differentiator between ranking markers, such as is suggested in the prior art, and using machine learning methods, as with the present invention, to model the combinations of these markers (i.e. modelling the space as a ‘fingerprint’). In practice, this means that the relative position of the markers on different patients might well vary, yet the method according to the invention still works at finding a correlation as a screening method and / or for predicting glucose levels.

[0054] As shown in Figure 2, methods according to the invention for generating a training dataset preferably utilise results from analysed breath samples, HbA1 tests, and oral glucose tolerance tests (OGTT).

[0055] According to a non-limiting example, a labelled training data set was generated, as follows:

[0056] Participants were each tested as follows:

[0057] 1 ) An HbA1 c test was performed.

[0058] 2) A standardised 75g oral glucose tolerance test (OGTT) was performed, on the following basis:

[0059] a. 8 hours of fasting in advance of the test.

[0060] b. A blood sample was taken (0 minutes).

[0061] c. 75g of glucose was subsequently administered orally.d. Further blood samples were taken at intervals of 30, 60, 90 and 120 minutes.

[0062] 3) Breath samples were taken at the time of the blood samples above, namely at 0, 30, 60, 90 and 120 minutes.

[0063] The results of the tests at (1 ) and (2), and the breath samples from (3), for each of the participants, were compiled in the dataset.

[0064] The union of compounds measured across all samples totalled 174 compounds (VOCs), as follows: Norflurane, Carbonyl sulfide, Dimethyl ether, Acetaldehyde, Butane, 1 -isocyano-, Ethanol, Isobutane, Acetonitrile, Methyl isocyanide, Acetone, Isopropyl Alcohol, Furan, Isoprene, Methylene chloride, Carbon disulfide, Acetic acid, n-Hexane, Furan, 2-methyl-, Tetrahydrofuran, Butanal, 3-methyl-, Benzene, 1 -Butanol, Butanal, 2-methyl-, Thiophene, 2-Propanol, 1-methoxy-, 2-Pentanone, Heptane, Pentanal, Furan, 2,5-dimethyl-, 1,4-Dioxane, 1 -Pentene, 2.4.4-trimethyl-, Butanenitrile, 3-methyl-, Methyl Isobutyl Ketone, Toluene, Thiophene, 3-methyl-, 1 -Octene, 2-Hexanone, Octane, Hexanal, Maleic anhydride, 1H-Pyrrole, 2,5-dimethyl-, 1 -Hexanol, Lactic acid, Benzene, 1 ,3-dimethyl-, Pentanoic acid, 1 -Nonene, 2-Heptanone, 4-Cyclopentene-1 ,3-dione, Styrene, Heptanal, Butyrolactone, Disulfide, methyl 2-propenyl, 2-Heptanone, 4-methyl-, 2,5-Furandione, dihydro-3-methylene-, 2(3H)-Furanone, dihydro-5-methyl-, Hexanoic acid, 1 -Nonanol, Benzaldehyde, Phenol, 5-Hepten-2-one, 6-methyl-, 2-Heptanone, 4,6-dimethyl-, Furan, 2-pentyl-, Benzonitrile, Decane, Hexanoic acid, ethyl ester, Octanal, Benzofuran, 1 -Hexanol, 2 -ethyl-, Cyclohexene, 1-methyl-4-(1-methylethenyl)-, (S)-, Benzeneacetaldehyde, Benzaldehyde, 2-hydroxy-, 2H-Pyran-2-one, tetrahydro-, 1 -Octanol, Acetophenone, Benzaldehyde, 2-methyl-, 4-Undecene, (E)-, Diallyl disulphide, Undecane, Nonanal, Benzofuran, 2-methyl-, 1-Propanol, 3,3'-oxybis-, Benzoic acid, Octanoic acid, Alkanoll, 1 ,2-Propanedione, 1-phenyl-, 1 -Dodecene, 2-Decanone, Dodecane, Ethanone, 1-(3-methylphenyl)-, Naphthalene, Decanal, 2-Propanol, 1-(2-butoxy-1 -methylethoxy)-, Benzothiazole, Nonanoic acid, 1 -Decanol, Tridecane, Undecanal, n-Decanoic acid, Pentanoic acid, 2.2.4-trimethyl-3-hydroxy- isobutyl ester, 1(3H)-lsobenzofuranone, Tetradecane, Dodecanal, Ethanone, 2, 2-dihydroxy-1 -phenyl-, 1 ,3-lsobenzofurandione, 4-methyl-, 5,9-Undecadien-2-one, 6,10-dimethyl-, (E)-, 4-Methylphthalic anhydride, Dodecane,1 -chloro-, Alkenel, 1 -Pentadecene, o-Cyanobenzoic acid, Pentadecane, 2,4-Di-tert-butylphenol, Tridecanal, 1-Penten-3-one, 1-(2,6,6-trimethyl-2-cyclohexen-1-yl)-, Stilbene, o-Hydroxybiphenyl, Lilial, Dodecanoic acid, Phenylmaleic anhydride, Pentadecane, 3-methyl-, Pentanoic acid, 2,2,4-trimethyl-3-carboxyisopropyl, isobutyl ester, Hexadecane, Phenol, 2-cyclohexyl-, Tetradecanal, 4-Cyclopentene-1 ,3-dione, 4-phenyl-, Sesquiterpenel , Benzophenone, Cyclopentaneacetic acid, 3-oxo-2-pentyl-, methyl ester, Heptadecane, o-Anisic acid, 4-nitrophenyl ester, Tetradecanoic acid, Alkyl phenol, Octanal, 2-(phenylmethylene)-, Ethanone, 1 ,2-diphenyl-, 9H-Fluoren-9-one, Octadecane, Hexadecane, 2,6, 10,14-tetramethyl-Benzenesulfonamide, N-butyl-, 2-Ethylhexyl salicylate, Isopropyl myristate, Benzene, 1,1'-[1,2-ethanediylbis(oxy)]bis-, Pentadecanoic acid, 10-Methylanthracene-9-carboxaldehyde, Oxacyclotetradecane-2,11 -dione, 13-methyl-, 7-Acetyl-6-ethyl-1 ,1 ,4,4-tetramethyltetralin, Cetene, Naphthalic anhydride, Nonadecane, Benzoic acid, 2-hydroxy-, phenylmethyl ester, (E,E)-7,11 ,15-Trimethyl-3-methylene-hexadeca-1 ,6,10,14-tetraene, Hexadecanoic acid, methyl ester, 9-Hexadecenoic acid, Nonadecane, 4-methyl-, n-Hexadecanoic acid, Diphenyl sulfone, Dibutyl phthalate, Eicosane, Isopropyl palmitate, Squalene, Hexanoic acid, 2-ethyl-hexadecyl ester, Eicosane, 3-methyl-, 1 -Octadecene, Methanone, 2-benzofuranylphenyl-, Heneicosane, 2, 3-Diphenylmaleic anhydride, Oleic Acid, Octadecanoic acid, Docosane, Benzoic acid, tetradecyl ester, p-Terphenyl, Tricosane, 2-methyl-, [1 ,1':3',1"-Terphenyl]-2'-ol, Tricosane, Hexacosane.

[0065] It is notable that the invention is not to be limited to the above compounds. In particular, the following additional compounds have further been identified: Ethane, Pentane, 2-Methylpentane, 3-Methylpentane, Formaldehyde, Propanal, Malondialdehyde, 4-Hydroxynonenal, 2-Butanone (methyl ethyl ketone), 3-Pentanone, Cyclohexanone, Acetoin, Diacetyl (2,3-butanedione), Methanol, 1-Propanol, 1 -Pentanol, Formic acid, Propionic acid, Butyric acid, Isovaleric acid, Methyl acetate, Ethyl acetate, Hydrogen sulfide, Methanethiol, Dimethyl sulfide, Dimethyl trisulfide, Ammonia, Nitric oxide, Nitrogen dioxide, Trimethylamine, Pyridine, Indole, Skatole, p-Cresol, o-Cresol, lndole-3-acetaldehyde, Limonene, Pinene, Geranyl acetone.The HbA1c and OGTT test results from 0, 30, 60 and 90 minutes provided for labelled breath sample data, with the diabetic status (healthy, pre-diabetic, and type 2 diabetes mellitus) and glucose levels of each participant known.

[0066] The glucose readings from the OGTT taken at times 0, 30, 60, and 90 minutes were used for training / mapping, whilst the glucose readings from the OGTT at 120 minutes were used for testing (to predict the glucose levels of the 120 minute sample for each participant).

[0067] Additional information for each participant was additionally used for training / mapping, specifically the body mass index (BMI), weight, height, sex and age of each participant were included alongside the compounds measured in the breath samples.

[0068] It should be appreciated that the above training data set generation is purely exemplary and should not be taken as limiting. Numerous alternatives will be possible.

[0069] For example, and without limitation:

[0070] Where only classification is of interest, i.e. glucose level prediction is not desired, the training data set may not include OGTT results. The training data set may include HbA1c results only.

[0071] Whether both OGTT and HbA1c results are used or not, the training data set may or may not include the additional information for each participant and / or the additional information may be varied. Moreover, the additional information need not include all of BMI, weight, height, sex and age.

[0072] Where OGTT results are used, the training data set may include more or less readings, in particular there may be only readings at 0 minutes and 120 minutes, with both readings input for training purposes.

[0073] Classification - healthy, prediabetic, Type 1 or Type 2 diabetes mellitus:The training set was used to create models for classifying input signals from the measured breath samples into their diabetes classification. Multiple models were used successfully, including K-nearest neighbour, K-means, Gaussian mixture models, SVMs and Convolutional Neural Networks. The most performant model, however, was a Random Forest Classifier, which, as will be appreciated by those skilled in the art, builds decision trees with several randomly sampled subsets of biomarkers. For each subset, it creates a tree that is used generate a determination from ‘Type T, ‘Type2’, ‘no diabetes’ or ‘pre-diabetic’. The decision trees are counted for making determinations. The Random Forest Classifier achieved an accuracy of 80% on the test dataset.

[0074] Regression - glucose level:

[0075] Multiple algorithms were used successfully, including linear models, random forests, and multi-layer perceptrons. The most performant model, however, with respect to R2was an Extra-Trees Regressor, which, as will be appreciated by those skilled in the art, is a meta estimator that fits a number of randomized decision trees on various sub-samples of the dataset and uses averaging to improve the predictive accuracy and control over-fitting. The Extra-Trees Regressor achieved an R2score of 0.755 on the test dataset.

[0076] The models may be trained, i.e. optimised, according to a single-objective loss function, using measures such as R2, MAE, Accuracy, or otherwise. An objective function may, however, be implemented that maps directly to specific real-world objectives, as will be readily appreciated by those skilled in the art.

[0077] By way of example only, consideration will now be given to specific determinations made, with reference to the flowchart of Figure 2.

[0078] Different relative combinations of the VOCs are determined. As should be appreciated different combinations will be noted from different patients, based for example, on their different ages, different food habits, and origination from / residence in different countries.Furthermore, different normalisations and different algorithms for regression and classification will model the breath in different ways, as again will be readily appreciated by those skilled in the art.

[0079] By way of example only, the top 10 markers for each of the 6 models developed for screening on the entire space for each of the three scenarios: 174 VOCs for all VOCs, 149 VOCs for the VOCs with saturated hydrocarbons (Sat HCs) removed, and 136 VOCs for VOCs with all hydrocarbons (HCs) removed are shown below in table 1. The remaining markers are excluded from the table below.

[0080] Each model was trained on the basis of the training data as detailed above and comprised a Random Forest Classifier. The normalisation applied was either Area or MaxSignal, as noted in table 1. As will be appreciated by those skilled in the art, MaxSignal relates to the peak height (signal intensity) and Area is the area under the peak normalised to the variation of the instrument. Area is more indicative of a concentration. Max Signal is more suited to qualitative work and peak detection.

[0081] Each model was able to successfully distinguish between the groups (healthy, pre, type 1 and type 2), yet the order in which some of the markers appeared changed, as demonstrated by the markers listed in the table, and some of the VOCs that might previously have been considered as ‘key markers’ e.g. isoprene, were not even present in the space used to model in the case of no hydrocarbons.

[0082]

[0083] Table 1 - Top 10 markers - classification models

[0084] As discussed, the order of some of the markers may change (which is a result 5 of the change in the space being modelled), as is illustrated by the presentation of the differing top 10 markers in table 1 , but that is not in and of itself what is used in making the final determination. Each model was accurate in successfully separating the groups.

[0085] io The positions of the specific VOCs are relative, and instead of ranking or assigning importance to specific markers, all VOCs are considered and the correlations between them, using appropriate machine learning models, such as the utilised random forest, to determine the number of decisions made for the combinedspace as an indicator of whether an individual is healthy, pre, type 1 or type 2

[0086] diabetes mellitus.

[0087] Highly accurate results are achieved.

[0088] 5

[0089] Considering the prediction of glucose levels, by way of example only, the top

[0090] 10 markers for each of an additional 6 models developed for screening on the entire space for each of the three scenarios: 174 VOCs for all VOCs, 149 VOCs for the VOCs with saturated hydrocarbons removed, and 136 VOCs for VOCs with all

[0091] io hydrocarbons removed) are shown below in table 2. The respective R2values are shown in table 3.

[0092] Each model was trained on the basis of the training data as detailed above

[0093] and comprised an Extra-Trees Regressor. The normalisation applied was again

[0094] 15 either Area or MaxSignal, as noted in table 2 below.

[0095] >

[0096]

[0097] Table 2 - Top 10 markers - regression models

[0098]

[0099] Table 3 - R2values - regression models

[0100] As will be apparent, there is a high correlation with measured blood glucose levels using different kinds of data normalisation and even different kinds of VOCs (all VOCs vs no HCs) with some like Squalene jumping from position 7 (no saturated HCs with r2 = 0.65) to position 22 (ALL VOCs with r2 = 0.68).

[0101] For the classification of T2 vs. others, there is a 93% accuracy with both Non-Saturated HCs and for the case of removing all HCs with the same normalisation, MaxSignal. Yet there is the case of o-Cyanobenzoic acid, which in the first case (no HCs MaxSignal) is found in position 7, and in the second case (no sat HCs MaxSignal), is found in position 102. This does not mean it is not important, rather, it remains part of the modelling space but has a different relative position. This demonstrates clearly that elements can change depending on normalisation and methods used, but that these elements analysed are important.

[0102] It follows that models can be readily developed, under the scope of the claimed invention, for both accurate classification and regression, which are not dependent on specific rankings of the VOCs.

[0103] Numerous alternative arrangements and modifications to the embodiments as described herein will be readily appreciated by those skilled in the art within the scope of the appended claims. In particular, features of the described embodiments may be readily combined.

[0104] It should be noted that whilst a number of models have been successfully implemented, numerous alternative discriminative or generative models used for regression / classification may also be implemented, as will be readily appreciated by those skilled in the art.When used in this specification and claims, the terms "comprises" and "comprising" and variations thereof mean that the specified features, steps or integers are included. The terms are not to be interpreted to exclude the presence of other features, steps or components.

[0105] The features disclosed in the foregoing description, or the following claims, or the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for attaining the disclosed result, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof.

[0106] Although certain example embodiments of the invention have been described, the scope of the appended claims is not intended to be limited solely to these embodiments. The claims are to be construed literally, purposively, and / or to encompass equivalents.

Claims

Claims1. A non-invasive diagnosis system, which comprises:a device for collecting a breath sample from a user, which comprises a mouthpiece;a spectrometer for measuring Volatile Organic Compounds, VOCs, in the breath sample; anda processor,wherein the processor is configured to analyse the measured VOCs using one or more machine learning models for classifying the user, based on one or more combinations of the measured VOCs, as healthy, prediabetic, Type 1 or Type 2 diabetes mellitus.

2. A non-invasive diagnosis system as claimed in Claim 1 , wherein the one or more machine learning models comprise a model generated using a labelled training data set, which comprises a plurality of data points, wherein each data point corresponds to an individual and is associated with one of the labels healthy, prediabetic, Type 1 or Type 2 diabetes mellitus for the individual, and each data point comprises the measured VOCs from one or more breath samples from the individual.

3. A non-invasive diagnosis system as claimed in Claim 2, wherein each data point is further associated with a glucose level of the individual.

4. A non-invasive diagnosis system as claimed in Claim 2 or 3, wherein each data point further comprises the result of an HbA1c test performed on the individual.

5. A non-invasive diagnosis system as claimed in any of Claims 2 to 4, wherein each data point further comprises the results of an oral glucose tolerance test, OGTT, performed on the individual.

6. A non-invasive diagnosis system as claimed in any of Claims 2 to 5, wherein each data point further comprises additional data on the individual, including one or more or all of age, sex, weight, height and body mass index.

7. A non-invasive diagnosis system as claimed in any of Claims 2 to 6, wherein the one or more machine learning models comprise a model generated using training data that comprises a plurality or all of the following VOCs: Norflurane, Carbonyl sulfide, Dimethyl ether, Acetaldehyde, Butane, 1 -isocyano-, Ethanol, Isobutane, Acetonitrile, Methyl isocyanide, Acetone, Isopropyl Alcohol, Furan, Isoprene, Methylene chloride, Carbon disulfide, Acetic acid, n-Hexane, Furan, 2-methyl-, Tetrahydrofuran, Butanal, 3-methyl-, Benzene, 1 -Butanol, Butanal, 2-methyl-, Thiophene, 2-Propanol, 1 -methoxy-, 2-Pentanone, Heptane, Pentanal, Furan, 2,5-dimethyl-, 1,4-Dioxane, 1 -Pentene, 2,4,4-trimethyl-, Butanenitrile, 3-methyl-, Methyl Isobutyl Ketone, Toluene, Thiophene, 3-methyl-, 1 -Octene, 2-Hexanone, Octane, Hexanal, Maleic anhydride, 1H-Pyrrole, 2,5-dimethyl-, 1-Hexanol, Lactic acid, Benzene, 1 ,3-dimethyl-, Pentanoic acid, 1 -Nonene, 2-Heptanone, 4-Cyclopentene-1 ,3-dione, Styrene, Heptanal, Butyrolactone, Disulfide, methyl 2-propenyl, 2-Heptanone, 4-methyl-, 2,5-Furandione, dihydro-3-methylene-2(3H)-Furanone, dihydro-5-methyl-, Hexanoic acid, 1 -Nonanol, Benzaldehyde, Phenol, 5-Hepten-2-one, 6-methyl-, 2-Heptanone, 4,6-dimethyl-, Furan, 2-pentyl-, Benzonitrile, Decane, Hexanoic acid, ethyl ester, Octanal, Benzofuran, 1 -Hexanol, 2-ethyl-, Cyclohexene, 1-methyl-4-(1 -methylethenyl)-, (S)-, Benzeneacetaldehyde, Benzaldehyde, 2-hydroxy-, 2H-Pyran-2-one, tetrahydro-, 1 -Octanol, Acetophenone, Benzaldehyde, 2-methyl-, 4-Undecene, (E)-, Diallyl disulphide, Undecane, Nonanal, Benzofuran, 2-methyl-, 1 -Propanol, 3,3'-oxybis-, Benzoic acid, Octanoic acid, Alkanoll, 1 ,2-Propanedione, 1 -phenyl-, 1 -Dodecene, 2-Decanone, Dodecane, Ethanone, 1-(3-methylphenyl)-, Naphthalene, Decanal, 2-Propanol, 1-(2-butoxy-1-methylethoxy)-, Benzothiazole, Nonanoic acid, 1 -Decanol, Tridecane, Undecanal, n-Decanoic acid, Pentanoic acid, 2,2,4-trimethyl-3-hydroxy-, isobutyl ester, 1(3H)-Isobenzofuranone, Tetradecane, Dodecanal, Ethanone, 2, 2-dihydroxy-1 -phenyl-, 1 ,3-lsobenzofurandione, 4-methyl-, 5,9-Undecadien-2-one, 6,10-dimethyl-, (E)-, 4-Methylphthalic anhydride, Dodecane, 1 -chloro-, Alkenel, 1 -Pentadecene, o-Cyanobenzoic acid, Pentadecane, 2,4-Di-tert-butylphenol, Tridecanal, 1-Penten-3-one, 1-(2,6,6-trimethyl-2-cyclohexen-1-yl)-, Stilbene, o-Hydroxybiphenyl, Lilial, Dodecanoic acid, Phenylmaleic anhydride, Pentadecane, 3-methyl-, Pentanoic acid, 2,2,4-trimethyl-3-carboxyisopropyl, isobutyl ester, Hexadecane, Phenol, 2-cyclohexyl-, Tetradecanal, 4-Cyclopentene-1 ,3-dione, 4-phenyl-, Sesquiterpenel ,Benzophenone, Cyclopentaneacetic acid, 3-oxo-2-pentyl-, methyl ester, Heptadecane, o-Anisic acid, 4-nitrophenyl ester, Tetradecanoic acid, Alkyl phenol, Octanal, 2-(phenylmethylene)-, Ethanone, 1 ,2-diphenyl-, 9H-Fluoren-9-one, Octadecane, Hexadecane, 2,6,10,14-tetramethyl-, Benzenesulfonamide, N-butyl-, 2-Ethylhexyl salicylate, Isopropyl myristate, Benzene, 1 , 1'-[1 ,2-ethanediylbis(oxy)]bis-, Pentadecanoic acid, 10-Methylanthracene-9-carboxaldehyde, Oxacyclotetradecane-2, 11 -dione, 13-methyl-, 7-Acetyl-6-ethyl-1,1,4,4-tetramethyltetralin, Cetene, Naphthalic anhydride, Nonadecane, Benzoic acid, 2-hydroxy-, phenylmethyl ester, (E, E)-7, 11 ,15-Trimethyl-3-methylene-hexadeca-1 ,6,10,14-tetraene, Hexadecanoic acid, methyl ester, 9-Hexadecenoic acid, Nonadecane, 4-methyl-, n-Hexadecanoic acid, Diphenyl sulfone, Dibutyl phthalate, Eicosane, Isopropyl palmitate, Squalene, Hexanoic acid, 2-ethyl-, hexadecyl ester, Eicosane, 3-methyl-, 1 -Octadecene, Methanone, 2-benzofuranylphenyl- Heneicosane, 2,3-Diphenylmaleic anhydride, Oleic Acid, Octadecanoic acid, Docosane, Benzoic acid, tetradecyl ester, p-Terphenyl, Tricosane, 2-methyl-, [1 ,1':3',1"-Terphenyl]-2'-ol, Tricosane, Hexacosane, Ethane, Pentane, 2-Methylpentane, 3-Methylpentane, Formaldehyde, Propanal, Malondialdehyde, 4-Hydroxynonenal, 2-Butanone (methyl ethyl ketone), 3-Pentanone, Cyclohexanone, Acetoin, Diacetyl (2,3-butanedione), Methanol, 1-Propanol, 1 -Pentanol, Formic acid, Propionic acid, Butyric acid, Isovaleric acid, Methyl acetate, Ethyl acetate, Hydrogen sulfide, Methanethiol, Dimethyl sulfide, Dimethyl trisulfide, Ammonia, Nitric oxide, Nitrogen dioxide, Trimethylamine, Pyridine, Indole, Skatole, p-Cresol, o-Cresol, lndole-3-acetaldehyde, Limonene, Pinene, Geranyl acetone.

8. A non-invasive diagnosis system as claimed in any of Claims 2 to 6, wherein the training data set excludes saturated hydrocarbons.

9. A non-invasive diagnosis system as claimed in any of Claims 2 to 6, wherein the one or more machine learning models comprise a model generated using training data that comprises a plurality or all of the following VOCs: Norflurane, Carbonyl sulfide, Dimethyl ether, Acetaldehyde, Butane, 1 -isocyano-, Ethanol, Acetonitrile, Methyl isocyanide, Acetone, Isopropyl Alcohol, Furan, Isoprene, Methylene chloride, Carbon disulfide, Acetic acid, Furan, 2-methyl-, Tetrahydrofuran, Butanal, 3-methyl-, Benzene, 1 -Butanol, Butanal, 2-methyl-, Thiophene, 2-Propanol,1 -methoxy-, 2-Pentanone, Pentanal, Furan, 2,5-dimethyl-, 1,4-Dioxane, 1 -Pentene, 2.4.4-trimethyl-, Butanenitrile, 3-methyl-, Methyl Isobutyl Ketone, Toluene, Thiophene, 3-methyl-, 1 -Octene, 2-Hexanone, Hexanal, Maleic anhydride, 1H-Pyrrole, 2,5-dimethyl-, 1 -Hexanol, Lactic acid, Benzene, 1 ,3-dimethyl-, Pentanoic acid, 1 -Nonene, 2-Heptanone, 4-Cyclopentene-1 ,3-dione, Styrene, Heptanal, Butyrolactone, Disulfide, methyl 2-propenyl, 2-Heptanone, 4-methyl-, 2,5-Furandione, dihydro-3-methylene-, 2(3H)-Furanone, dihydro-5-methyl-, Hexanoic acid, 1 -Nonanol, Benzaldehyde, Phenol, 5-Hepten-2-one, 6-methyl-, 2-Heptanone, 4,6-dimethyl-, Furan, 2-pentyl-, Benzonitrile, Hexanoic acid, ethyl ester, Octanal, Benzofuran, 1 -Hexanol, 2-ethyl-, Cyclohexene, 1-methyl-4-(1 -methylethenyl)-, (S)-, Benzeneacetaldehyde, Benzaldehyde, 2-hydroxy-, 2H-Pyran-2-one, tetrahydro-, 1-Octanol, Acetophenone, Benzaldehyde, 2-methyl-, 4-Undecene, (E)-, Diallyl disulphide, Nonanal, Benzofuran, 2-methyl-, 1 -Propanol, 3,3'-oxybis-, Benzoic acid, Octanoic acid, Alkanoll, 1,2-Propanedione, 1 -phenyl-, 1 -Dodecene, 2-Decanone, Ethanone, 1-(3-methylphenyl)-, Naphthalene, Decanal, 2-Propanol, 1-(2-butoxy-1-methylethoxy)-, Benzothiazole, Nonanoic acid, 1 -Decanol, Undecanal, n-Decanoic acid, Pentanoic acid, 2,2,4-trimethyl-3-hydroxy-, isobutyl ester, 1(3H)-Isobenzofuranone, Dodecanal, Ethanone, 2, 2-dihydroxy-1 -phenyl-, 1,3-Isobenzofurandione, 4-methyl-, 5,9-Undecadien-2-one, 6,10-dimethyl-, (E)-, 4-Methylphthalic anhydride, Dodecane, 1 -chloro-, Alkenel, 1 -Pentadecene, o-Cyanobenzoic acid, 2,4-Di-tert-butylphenol, Tridecanal, 1-Penten-3-one, 1 -(2,6,6-trimethyl-2-cyclohexen-1-yl)-, Stilbene, o-Hydroxybiphenyl, Lilial, Dodecanoic acid, Phenylmaleic anhydride, Pentanoic acid, 2,2,4-trimethyl-3-carboxyisopropyl, isobutyl ester, Phenol, 2-cyclohexyl-, Tetradecanal, 4-Cyclopentene-1 ,3-dione, 4-phenyl-, Sesquiterpenel , Benzophenone, Cyclopentaneacetic acid, 3-oxo-2-pentyl-, methyl ester, o-Anisic acid, 4-nitrophenyl ester, Tetradecanoic acid, Alkyl phenol, Octanal, 2-(phenylmethylene)-, Ethanone, 1 ,2-diphenyl-, 9H-Fluoren-9-one, Benzenesulfonamide, N-butyl-, 2-Ethylhexyl salicylate, Isopropyl myristate, Benzene, 1,1'-[1,2-ethanediylbis(oxy)]bis-, Pentadecanoic acid, 10-Methylanthracene-9-carboxaldehyde, Oxacyclotetradecane-2,11 -dione, 13-methyl-, 7-Acetyl-6-ethyl- 1.1.4.4-tetramethyltetralin, Cetene, Naphthalic anhydride, Benzoic acid, 2-hydroxy-, phenylmethyl ester, (E, E)-7, 11 ,15-Trimethyl-3-methylene-hexadeca-1 ,6,10,14-tetraene, Hexadecanoic acid, methyl ester, 9-Hexadecenoic acid, n-Hexadecanoic acid, Diphenyl sulfone, Isopropyl palmitate, Squalene, Hexanoic acid, 2-ethyl-,hexadecyl ester, 1 -Octadecene, Methanone, 2-benzofuranylphenyl-, 2,3-Diphenylmaleic anhydride, Oleic Acid, Octadecanoic acid, Benzoic acid, tetradecyl ester, p-Terphenyl, [1 ,1':3',1"-Terphenyl]-2'-ol, Formaldehyde, Propanal, Malondialdehyde, 4-Hydroxynonenal, 2-Butanone (methyl ethyl ketone), 3-Pentanone, Cyclohexanone, Acetoin, Diacetyl (2,3-butanedione), Methanol, 1-Propanol, 1 -Pentanol, Formic acid, Propionic acid, Butyric acid, Isovaleric acid, Methyl acetate, Ethyl acetate, Hydrogen sulfide, Methanethiol, Dimethyl sulfide, Dimethyl trisulfide, Ammonia, Nitric oxide, Nitrogen dioxide, Trimethylamine, Pyridine, Indole, Skatole, p-Cresol, o-Cresol, lndole-3-acetaldehyde, Limonene, Pinene, Geranyl acetone.

10. A non-invasive diagnosis system as claimed in any of Claims 2 to 6, wherein the training data set excludes all hydrocarbons.

11. A non-invasive diagnosis system as claimed in any of Claims 2 to 6, wherein the one or more machine learning models comprise a model generated using training data that comprises a plurality or all of the following VOCs: Norflurane, Carbonyl sulfide, Dimethyl ether, Acetaldehyde, Butane, 1 -isocyano-, Ethanol, Acetonitrile, Methyl isocyanide, Acetone, Isopropyl Alcohol, Furan, Methylene chloride, Carbon disulfide, Acetic acid, Furan, 2-methyl-, Tetrahydrofuran, Butanal, 3-methyl-, Benzene, 1 -Butanol, Butanal, 2-methyl-, Thiophene, 2-Propanol, 1-methoxy-, 2-Pentanone, Pentanal, Furan, 2,5-dimethyl-, 1,4-Dioxane, Butanenitrile, 3-methyl-, Methyl Isobutyl Ketone, Toluene, Thiophene, 3-methyl-, 2-Hexanone, Hexanal, Maleic anhydride, 1H-Pyrrole, 2,5-dimethyl-, 1 -Hexanol, Lactic acid, Benzene, 1 ,3-dimethyl-, Pentanoic acid, 2-Heptanone, 4-Cyclopentene-1 ,3-dione, Styrene, Heptanal, Butyrolactone, Disulfide, methyl 2-propenyl, 2-Heptanone, 4-methyl-, 2,5-Furandione, dihydro-3-methylene-, 2(3H)-Furanone, dihydro-5-methyl-, Hexanoic acid, 1 -Nonanol, Benzaldehyde, Phenol, 5-Hepten-2-one, 6-methyl-, 2-Heptanone, 4,6-dimethyl-, Furan, 2-pentyl-, Benzonitrile, Hexanoic acid, ethyl ester, Octanal, Benzofuran, 1 -Hexanol, 2 -ethyl-, Benzeneacetaldehyde, Benzaldehyde, 2-hydroxy-, 2H-Pyran-2-one, tetrahydro-, 1 -Octanol, Acetophenone, Benzaldehyde, 2-methyl-, Diallyl disulphide, Nonanal, Benzofuran, 2-methyl-, 1 -Propanol, 3,3'-oxybis-, Benzoic acid, Octanoic acid, Alkanoll, 1,2-Propanedione, 1 -phenyl-, 2-Decanone, Ethanone, 1-(3-methylphenyl)-, Naphthalene, Decanal, 2-Propanol, 1-(2-butoxy-1-methylethoxy)-, Benzothiazole, Nonanoic acid, 1 -Decanol, Undecanal, n-Decanoic acid, Pentanoic acid, 2,2,4-trimethyl-3-hydroxy-, isobutyl ester, 1(3H)-Isobenzofuranone, Dodecanal, Ethanone, 2, 2-dihydroxy-1 -phenyl-, 1,3-Isobenzofurandione, 4-methyl-, 5,9-Undecadien-2-one, 6,10-dimethyl-, (E)-, 4-Methylphthalic anhydride, Dodecane, 1 -chloro-, o-Cyanobenzoic acid, 2,4-Di-tert-butylphenol, Tridecanal, 1 -Penten-3-one, 1-(2,6,6-trimethyl-2-cyclohexen-1-yl)-, Stilbene, o-Hydroxybiphenyl, Lilial, Dodecanoic acid, Phenylmaleic anhydride, Pentanoic acid, 2,2,4-trimethyl-3-carboxyisopropyl, isobutyl ester, Phenol, 2-cyclohexyl-, Tetradecanal, 4-Cyclopentene-1 ,3-dione, 4-phenyl-, Benzophenone, Cyclopentaneacetic acid, 3-oxo-2-pentyl-, methyl ester, o-Anisic acid, 4-nitrophenyl ester, Tetradecanoic acid, Alkyl phenol, Octanal, 2-(phenylmethylene)-, Ethanone, 1 ,2-diphenyl-, 9H-Fluoren-9-one, Benzenesulfonamide, N-butyl-, 2-Ethylhexyl salicylate, Isopropyl myristate, Benzene, 1,1'-[1,2-ethanediylbis(oxy)]bis-Pentadecanoic acid, 10-Methylanthracene-9-carboxaldehyde, Oxacyclotetradecane-2, 11 -dione, 13-methyl-, 7-Acetyl-6-ethyl-1,1,4,4-tetramethyltetralin, Cetene, Naphthalic anhydride, Benzoic acid, 2-hydroxy-, phenylmethyl ester, Hexadecanoic acid, methyl ester, 9-Hexadecenoic acid, n-Hexadecanoic acid, Diphenyl sulfone, Isopropyl palmitate, Hexanoic acid, 2 -ethyl-, hexadecyl ester, Methanone, 2-benzofuranylphenyl-, 2,3-Diphenylmaleic anhydride, Oleic Acid, Octadecanoic acid, Benzoic acid, tetradecyl ester, p-Terphenyl, [1 ,1':3',1"-Terphenyl]-2'-ol, Formaldehyde, Propanal, Malondialdehyde, 4-Hydroxynonenal, 2-Butanone (methyl ethyl ketone), 3-Pentanone, Cyclohexanone, Acetoin, Diacetyl (2,3-butanedione), Methanol, 1 -Propanol, 1 -Pentanol, Formic acid, Propionic acid, Butyric acid, Isovaleric acid, Methyl acetate, Ethyl acetate, Hydrogen sulfide, Methanethiol, Dimethyl sulfide, Dimethyl trisulfide, Ammonia, Nitric oxide, Nitrogen dioxide, Trimethylamine, Pyridine, Indole, Skatole, p-Cresol, o-Cresol, lndole-3-acetaldehyde, Geranyl acetone.

12. A non-invasive diagnosis system as claimed in any of Claims 2 to 11 , wherein the training data set includes at least 10 VOCs, preferably at least 20 VOCs, and more preferably at least 30 VOCs.

13. A non-invasive diagnosis system as claimed in any preceding claim, wherein the one or more machine learning models comprise one or more discriminative or generative machine learning models.

14. A non-invasive diagnosis system as claimed in any preceding claim, wherein the one or more machine learning models are selected from: K-nearest neighbour, K-means, Gaussian Mixture Model, Support Vector Machine, Convolutional Neural Network, Random Forest Classifier, multi-layer perceptron, linear model, Extra Trees Regressor.

15. A method for classifying individuals using machine learning, the method comprising:training a machine learning model using a labelled training dataset, which comprises a plurality of data points, wherein each data point corresponds to an individual and is associated with one of the labels healthy, prediabetic, Type 1 or Type 2 diabetes mellitus for the individual, and each data point comprises the measured VOCs from one or more breath samples from the individual;applying the trained machine learning model to classify an individual based on one or more breath samples taken from the individual; andproviding the classification result.

16. The method of Claim 15 further comprising updating or augmenting the training dataset with real or synthetic data.

17. The method of Claim 15 or 16 further comprising improving or retraining the machine learning model based on user-provided corrections or newly labelled data.