Biomarkers
Patent Information
- Application Number
- JP2024503786
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-19
- Filing Date
- 2022-07-19
- Publication Date
- 2025-07-29
AI Technical Summary
Current diagnostic methods for cardiopulmonary diseases causing shortness of breath, such as asthma, COPD, heart failure, and pneumonia, are inadequate due to low discriminatory power and delays in blood sample processing, leading to inappropriate treatment decisions and adverse patient outcomes.
The use of volatile organic compound (VOC) biomarkers detected in exhaled breath samples through methods like two-dimensional gas chromatography combined with mass spectrometry to diagnose and treat cardiopulmonary diseases, allowing for non-invasive and patient-specific diagnosis and treatment.
The VOC biomarkers demonstrate high sensitivity and specificity, with up to 79% sensitivity and 85% specificity in differentiating individuals with cardiopulmonary diseases from healthy individuals, enabling informed clinical decisions and effective treatment selection.
Smart Images

Figure 00000041_0000 
Figure 00000041_0001 
Figure 00000041_0002
Abstract
Description
[Technical field]
[0001] FIELD OF THEINVENTION The present invention relates to methods of diagnosing and treating cardiorespiratory disease in subjects experiencing shortness of breath. [Background technology]
[0002] background Shortness of breath due to cardiopulmonary disease accounts for more than one in eight of all emergency admissions. Despite the same symptom presentation, the etiology of acute breathlessness is highly diverse, as are the disease courses and treatment options. Diagnostic evaluation of acute breathlessness relies heavily on investigations such as blood-based biomarkers (e.g., C-reactive protein (CRP), B-type natriuretic peptide (NT-proBNP)) and radiological tests. Although these biomarkers have clinical utility for patients with mostly single pathologies, they have poor discrimination power for patients with multifactorial symptoms of acute breathlessness and are particularly difficult to interpret in the context of pre-hospital treatment exposures (e.g., antibiotics for pneumonia and CRP values at admission). Furthermore, delays in blood sample processing at the point of triage can result in inappropriate treatment decisions and, consequently, adversely affect patients. To address these issues, there have been significant advances in the field of metabolomics, supported by analytical techniques, which have enabled the comprehensive identification and quantification of metabolite profiles in biological systems from samples taken at the point of clinical care.
[0003] Nevertheless, there is a need for biomarkers that can be used to diagnose and differentiate cardiopulmonary diseases manifested by shortness of breath. Summary of the Invention
[0004] Description of the invention According to a first aspect of the present invention there is provided a method of diagnosing cardiopulmonary disease in a subject, the method comprising: detecting the presence of one or more cardiopulmonary disease-VOC biomarkers in a sample of exhaled breath from the subject; Here, if one or more VOC biomarkers are present in the sample, the subject may be suffering from cardiopulmonary disease.
[0005] In one embodiment, a method of diagnosing asthma in a subject is provided, the method comprising: detecting the presence of one or more asthma-VOC biomarkers in an exhaled breath sample from the subject; Here, if one or more VOC biomarkers are present in the sample, the subject may suffer from asthma.
[0006] In one embodiment, a method of diagnosing COPD in a subject is provided, the method comprising: detecting the presence of one or more COPD-VOC biomarkers in an exhaled breath sample from the subject; Wherein, if one or more VOC biomarkers are present in the sample, the subject may be suffering from COPD.
[0007] In one embodiment, a method of diagnosing pneumonia in a subject is provided, the method comprising: detecting the presence of one or more pneumonia-VOC biomarkers in an exhaled breath sample from the subject; Here, if one or more VOC biomarkers are present in the sample, the subject may be suffering from pneumonia.
[0008] In one embodiment, a method of diagnosing heart failure in a subject is provided, the method comprising: detecting the presence of one or more heart failure VOC biomarkers in an exhaled breath sample from a subject; Here, if one or more VOC biomarkers are present in the sample, the subject may be suffering from heart failure.
[0009] According to a second aspect, there is provided a method of treating cardiopulmonary disease in a subject, the method comprising: detecting the presence of one or more cardiopulmonary disease-VOC biomarkers in an exhaled breath sample from the subject; wherein the presence of one or more VOC biomarkers in the sample indicates that the subject is suffering from cardiopulmonary disease; and This involves administering a therapeutic agent to the subject to treat the cardiopulmonary disease.
[0010] According to a third aspect, there is provided a method of treating cardiopulmonary disease in a subject, the method comprising: The method includes administering a therapeutic agent to a subject diagnosed with cardiopulmonary disease using a method according to the invention.
[0011] According to a fourth aspect, there is provided a method of selecting a subject for treatment with a cardiorespiratory disease therapeutic agent or composition, the method comprising: detecting the presence of one or more cardiopulmonary disease-VOC biomarkers in an exhaled breath sample from the subject; wherein the presence of one or more VOC biomarkers in the sample indicates that the subject is suffering from cardiopulmonary disease; and Selecting the subject for treatment with a cardiopulmonary disease therapeutic agent or composition.
[0012] According to another aspect, a method of selecting a subject for treatment with a cardiorespiratory disease therapeutic or composition is provided, the method comprising: Selecting a subject diagnosed with cardiopulmonary disease using a method according to the invention for treatment with a cardiopulmonary disease therapeutic agent or composition.
[0013] The present invention provides a more patient-based method for diagnosing and treating cardiopulmonary diseases. The present invention allows for the diagnosis of a subject without the use of invasive procedures such as blood sampling or radiological processes. The method does not have to be performed on a subject.
[0014] Two important characteristics of biomarkers used for diagnostic purposes are sensitivity and specificity. The higher the degree of sensitivity, the lower the probability of false negatives. The higher the specificity, the lower the probability of false positives. The biomarkers disclosed herein can surprisingly show up to 79% sensitivity and 85% specificity (with AUC0.89) when distinguishing between individuals with cardiopulmonary disease and healthy individuals (controls).
[0015] The values for distinguishing each acute cardiopulmonary disease group from other acute cardiopulmonary disease groups (i.e., not for healthy patients) were as follows: · Asthma - sensitivity 0.75 (0.63, 0.85), specificity 0.90 (0.85, 0.94); · COPD-sensitivity 0.66 (0.52, 0.78), specificity 0.89 (0.85, 0.93); Heart failure - sensitivity 0.64 (0.48, 0.78), specificity 0.96 (0.92, 0.98); and Pneumonia - sensitivity 0.65 (0.51, 0.78), specificity 0.93 (0.89, 0.96).
[0016] Thus, the present invention enables clinicians to make more informed decisions regarding the diagnosis and treatment of subjects experiencing shortness of breath and suffering from cardiopulmonary disorders.
[0017] According to a fifth aspect, there is provided a method of determining whether a therapeutic agent or composition is effectively treating a cardiopulmonary disease in a subject, the method comprising: Measuring the concentration of one or more cardiopulmonary VOC biomarkers in a test sample exhaled by the subject; and comparing the concentration of at least one or more VOCs in a test sample to the concentration in a reference sample; Here, a lower concentration of one or more VOC biomarkers in the test sample compared to the concentration in the reference sample indicates that the therapeutic agent or composition is effectively treating the cardiopulmonary disease in the subject.
[0018] It will be understood that the concentration of the VOC biomarker in the test sample is positively correlated with the magnitude / severity of the cardiopulmonary disease. Thus, for example, a decrease in the concentration of the VOC biomarker in the test sample compared to the concentration in the reference sample may indicate a decrease in the magnitude / severity of the cardiopulmonary disease. Similarly, an increase in the concentration of the VOC biomarker in the test sample compared to the concentration in the reference sample may indicate an increase in the magnitude / severity of the cardiopulmonary disease.
[0019] The concentration of the VOC biomarker in the test sample may be at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% lower (or at least reduced) compared to the concentration in the reference sample.
[0020] The reference sample may be taken from the same subject or from a different subject. Preferably, the reference sample is taken from the same subject but at an earlier time point than the test sample. Preferably, the earlier sample indicates that the subject suffers from cardiopulmonary disease.
[0021] Preferably, the subject referred to herein is experiencing shortness of breath. Preferably, two-dimensional gas chromatography combined with mass spectrometry is used to detect the presence of one or more VOC biomarkers in the sample.
[0022] Cardiopulmonary disease may be a disease or disorder of the cardiovascular system and / or a disease of the respiratory system.Examples of cardiopulmonary disease include asthma, COPD, heart failure and respiratory infection (such as pneumonia), bronchitis, emphysema, congestive heart failure, hypertension, angina pectoris, peripheral vascular disease and myocardial infarction.Preferably, the term "cardiopulmonary disease" refers to one or more diseases selected from the group including asthma, COPD, heart failure and pneumonia.
[0023] VOCs (volatile organic compounds) may be referred to as organic compounds having a boiling point between about 50° C. and about 250° C. at standard atmospheric pressure of 101.3 kPa.
[0024] The cardiopulmonary disease-VOC biomarkers may be one or more selected from the group including hydrocarbons, ketones, aldehydes, alcohols, oxygen-containing VOCs, terpenoids, aromatics, sulfur-containing VOCs, nitrogen-containing VOCs, halogen salts (e.g., dichloromethane), and surfactants and emollients.
[0025] It will be appreciated that the step of detecting the presence of one or more cardiopulmonary disease-VOC biomarkers may comprise using a method according to the invention. It will also be appreciated that the detection of another cardiopulmonary disease-VOC biomarker in a sample is indicative that the subject (from whom the sample was taken) suffers from cardiovascular disease.
[0026] The hydrocarbon VOC may be one or more selected from the group including: 2-methylbutane; isoprene; 3-methylpentane; 2,4-dimethylpentane; 2,2-dimethylpentane; hexane; octane; 2,6-dimethyloctane; nonane; 2-methylnonane; 5-methylnonane; decane; 4-methyldecane; undecane; 4-methylundecane; dimethylundecane isomers; 3-methyltridecane; tetradecane; octadecane; 1-nonene; 1-decene; cyclohexane; cyclohexadiene isomers; methylcyclopentadiene; and hexadecene isomers. The hydrocarbon may be one or more selected from FIG. 16. Preferably, the hydrocarbon VOC may be one or more selected from the group including: hexane, octane, 2,6-dimethyloctane, nonane, 2-methylnonane, decane, undecane, 4-methylundecane, dimethylundecane isomers, 3-methyltridecane, tetradecane, octadecane, 1-nonene, 1-decene, cyclohexane, cyclohexadiene isomers, methylcyclopentadiene, and hexadecene isomers. The hydrocarbon may be one or more selected from FIG. 16.
[0027] The ketone VOC may be one or more selected from the group including: acetone; 2,3-butanedione; 2-pentanone; 3-buten-2-one (methyl vinyl ketone); 4-methyl-2-pentanone; 6-methyl-5-hepten-2-one; and cyclohexanone. The ketone may be one or more selected from Figure 16. Preferably, the ketone VOC may be one or more selected from the group including: 3-buten-2-one (methyl vinyl ketone); 4-methyl-2-pentanone; and 6-methyl-5-hepten-2-one.
[0028] The aldehyde VOC may be one or more selected from the group including: butanal; hexanal; nonanal; decanal; methyl decanal isomers; undecanal; 2-methyl-2-propenal (methacrolein); 3-methyl-benzaldehyde; and tridecanal. The aldehyde may be one or more selected from Figure 16. Preferably, the aldehyde VOC is one or more selected from the group including butanal; methyl decanal isomers; undecanal; 2-methyl-2-propenal (methacrolein); 3-methyl-benzaldehyde; and tridecanal.
[0029] The alcohol VOC may be one or more selected from the group including 2-propanol; 2-ethylhexanol; 1-decanol; and 1-hexadecanol. The alcohol may be one or more selected from FIG.
[0030] The oxygen-containing VOCs may be one or more selected from the group including: ethyl acetate; tetrahydrofuran; 1,4-dioxane; 2-methyl-1,3-dioxolane; and 1,3-dioxolane. The oxygen-containing VOCs may be one or more selected from FIG.
[0031] The terpenoid VOC may be one or more selected from the group including: limonene, alpha-pinene, eucalyptol, menthone, menthol, camphene, p-mentha-1,4 / 8-diene, 3-carene, beta-myrcene, beta-phellandrene, geranylacetone, beta-bisabolene, alpha-isomethylionone, and galaxolide. The terpenoid may be one or more selected from FIG. 16.
[0032] The aromatic VOC may be one or more selected from the group including: ethyl-benzene; 2,3-dimethylnaphthalene; and substituted benzene. The aromatic may be one or more selected from FIG.
[0033] The sulfur-containing VOC may be one or more selected from the group including 3-methylthiophene; dimethyl sulfide; allylmethyl sulfide; carbonyl sulfide; 1-(methylthio)-1-propene; and 1-methylthiopropane. The sulfur-containing VOC may be one or more selected from FIG. 16. Preferably, the sulfur-containing VOC may be one or more selected from the group including dimethyl sulfide; 1-(methylthio)-1-propene; and 1-methylthiopropane.
[0034] The nitrogen-containing VOCs may be one or more selected from the group including: 4-cyanocyclohexene; and methenamine. The nitrogen-containing VOCs may be one or more selected from FIG.
[0035] The surfactant and emollient VOC may be one or more selected from the group including: isopropyl myristate; stearyl vinyl ether; N,N-dimethyl-1-nonamine; N,N-dimethyl-1-dodecanamine; alkenyl hexanoate; 2,2,4,4,6,8,8-heptamethylnonane; dodecyl acrylate; and decyl isobutyl ether. The surfactant and emollient may be one or more selected from FIG. 16.
[0036] The cardiopulmonary disease-VOC biomarkers may be any combination of the VOC biomarkers disclosed in Figure 16. Thus, the cardiopulmonary disease-VOC biomarkers may be one or more selected from Figure 16. The VOCs may be isomers of the VOCs disclosed in Figure 16. Thus, the cardiopulmonary disease-VOC biomarkers may be one or more VOC biomarkers selected from Figure 16, or isomers thereof. The isomers may be structural isomers, diastereomers (e.g., cis-trans isomers or rotamers) or enantiomers.
[0037] In one embodiment, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose cardiopulmonary disease in a subject: hexane; octane; tetradecane, 2,3-butanedione; hexanal; 2-methyl-2-propenal; 1-hexadecanol; 2-methyl-1,3-dioxolane; limonene; eucalyptol; menthone; p-mentha-1,4 / 8-diene; 3-carene; beta-phellandrene; sesquiterpenoids; xylene; 2,3-dimethylnaphthalene; carbonyl sulfide; 4-cyanocyclohexene; methenamine; dichloromethane; N,N-dimethyl-1-nonamine; and alkenylhexanoate esters.
[0038] In another embodiment, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose asthma in a subject: hexane; 2-methylnonane; decane; tetradecane; 1-nonene; 2,3-butanedione; 2-pentanone; hexanal; nonanal; decanal; methyldecanal isomers; undecanal; 3-methyl-benzaldehyde; 2-ethylhexanol; 1-hexadecanol; tetrahydrofuran; 1,4-dioxane; 2-methyl-1,3-dioxolane; eucalyptol; p-mentha-1,4 / 8-diene; 3-carene; beta-phellandrene; β-bisabolene; sesquiterpenoids; xylene; 4-cyanocyclohexene; methenamine; stearyl vinyl ether; N,N-dimethyl-1-nonanamine; and N,N-dimethyl-1-dodecanamine.
[0039] A selection of one or more (e.g., all) of the following VOC biomarkers may be used to diagnose asthma in a subject: 3-methylpentane; hexane; 2-methylnonane; decane; tetradecane; 1-nonene; 2,3-butanedione; methyldecanal isomers; undecanal; 3-methyl-benzaldehyde; 2-ethylhexanol; 1-hexadecanol; tetrahydrofuran; 1,4-dioxane; 2-methyl-1,3-dioxolane; eucalyptol; p-mentha-1,4 / 8-diene; 3-carene; beta-phellandrene; β-bisabolene; sesquiterpenoids; xylene; 4-cyanocyclohexene; methenamine; stearyl vinyl ether; N,N-dimethyl-1-nonanamine; and N,N-dimethyl-1-dodecanamine.
[0040] Preferably, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose asthma in a subject: 2-methylnonane, decane, 1-nonene, 2-pentanone, nonanal, decanal, methyldecanal isomers, undecanal, 3-methyl-benzaldehyde, 2-ethylhexanol, tetrahydrofuran, 1,4-dioxane, β-bisabolene, and N,N-dimethyl-1-dodecanamine.
[0041] Most preferably, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose asthma in a subject: 2-methylnonane; decane; 1-nonene; methyldecanal isomers; undecanal; 3-methyl-benzaldehyde; 2-ethylhexanol; tetrahydrofuran; 1,4-dioxane; β-bisabolene; and N,N-dimethyl-1-dodecanamine.
[0042] In another embodiment, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose COPD in a subject: nonane; 4-methylundecane; cyclohexane; methylcyclopentadiene; 2,3-butanedione; 6-methyl-5-hepten-2-one; 1-decanol; eucalyptol; 2-methyl-1,3-dioxolane; limonene; menthol; camphene; menthone; galaxolide; 2,3-dimethylnaphthalene; carbonyl sulfide; 3-methylthiophene; alkenylhexanoic acid esters; allylmethyl sulfide; dichloromethane; and N,N-dimethyl-1-dodecanamine.
[0043] A selection of one or more (e.g., all) of the following VOC biomarkers may be used to diagnose COPD in a subject: nonane; 4-methylundecane; cyclohexane; methylcyclopentadiene; 6-methyl-5-hepten-2-one; 1-decanol; eucalyptol; 2-methyl-1,3-dioxolane; limonene; menthol; camphene; menthone; galaxolide; 2,3-dimethylnaphthalene; 3-methylthiophene; alkenylhexanoate; dichloromethane; and N,N-dimethyl-1-dodecanamine.
[0044] Preferably, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose COPD in a subject: 4-methylundecane; 1-decanol; menthol; camphene; galaxolide; 3-methylthiophene; and N,N-dimethyl-1-dodecanamine.
[0045] In another embodiment, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose heart failure in a subject: hexane; 5-methylnonane; 4-methyldecane; undecane; cyclohexene; acetone; butanal; 2-methyl-2-propenal; tridecanal; ethyl acetate; 1,3-dioxolane; limonene; 3-carene; beta-myrcene; ethyl-benzene; 2,3-dimethylnaphthalene; N,N-dimethyl-1-nonamine; 2-methyl-2-propenal (methacrolein); alkenyl hexanoic acid esters; and decyl isobutyl ether.
[0046] A selection of one or more (e.g., all) of the following VOC biomarkers may be used to diagnose heart failure in a subject: hexane; undecane; cyclohexene; acetone; butanal; 2-methyl-2-propenal; tridecanal; ethyl acetate; 1,3-dioxolane; limonene; 3-carene; beta-myrcene; ethyl-benzene; 2,3-dimethylnaphthalene; N,N-dimethyl-1-nonamine; 2-methyl-2-propenal (methacrolein); alkenyl hexanoate; and decyl isobutyl ether.
[0047] Preferably, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose heart failure in a subject: 5-methylnonane; 4-methyldecane; undecane; cyclohexene; butanal; 2-methyl-2-propenal; tridecanal; ethyl acetate; 1,3-dioxolane; beta-myrcene; ethyl-benzene; and decyl isobutyl ether.
[0048] Most preferably, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose heart failure in a subject: cyclohexene; butanal; 2-methyl-2-propenal; tridecanal; ethyl acetate; 1,3-dioxolane; beta-myrcene; ethyl-benzene; and decyl isobutyl ether.
[0049] In another embodiment, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose pneumonia in a subject: 2,4-dimethylpentane; 2,2-dimethylpentane; hexane; octane number; 2,6-dimethyloctane; dimethylundecane isomers; tetradecane; p-mentha-1,4 / 8-diene; 1-decene; 3-buten-2-one (methyl vinyl ketone); cyclohexanone; hexanal; 2-methyl-2-propenal; 2-propanol; 1-hexadecanol; α-pinene; menthone; beta phellandrene; sesquiterpenoids; xylene; carbonyl sulfide; 1-(methylthio)-1-propene; 2-methyl-2-propenal (methacrolein); 1-methylthiopropane; 4-cyanocyclohexene; methenamine; dichloromethane; and dodecyl acrylate.
[0050] A selection of one or more (e.g., all) of the following VOC biomarkers may be used to diagnose pneumonia in a subject: hexane; octane number; 2,6-dimethyloctane; dimethylundecane isomers; tetradecane; p-mentha-1,4 / 8-diene; 1-decene; 3-buten-2-one (methyl vinyl ketone); hexanal; 2-methyl-2-propenal; 2-propanol; 1-hexadecanol; α-pinene; menthone; beta-phellandrene; sesquiterpenoids; xylene; carbonyl sulfide; 1-(methylthio)-1-propene; 2-methyl-2-propenal (methacrolein); 1-methylthiopropane; 4-cyanocyclohexene; methenamine; dichloromethane; and dodecyl acrylate.
[0051] Preferably, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose pneumonia in a subject: 2,4-dimethylpentane; 2,2-dimethylpentane; 2,6-dimethyloctane; dimethylundecane isomers; 1-decene; 3-buten-2-one (methyl vinyl ketone); 1-(methylthio)-1-propene; 1-methylthiopropane; and dodecyl acrylate.
[0052] Most preferably, a selection of one or more (e.g., all) of the following VOC biomarkers are used to diagnose pneumonia in a subject: dimethylundecane isomers; 1-decene; 3-buten-2-one (methyl vinyl ketone); 1-(methylthio)-1-propene; 1-methylthiopropane; and dodecyl acrylate.
[0053] One or more of the cardiopulmonary disease-VOC biomarkers disclosed herein (e.g., an asthma-VOC biomarker, a COPD-VOC biomarker, a pneumonia-VOC biomarker, or a heart failure-VOC biomarker) may be a selection of one or more of the biomarkers disclosed above for diagnosing cardiopulmonary disease in a subject.
[0054] Preferably, one or more (eg, all) of the following VOC biomarkers are used to diagnose cardiopulmonary disease in a subject experiencing shortness of breath.
[0055] Detection of a single VOC biomarker may be used to diagnose a cardiopulmonary disease in a subject. However, it will be understood that the more VOC biomarkers used in the present invention, the more certain the diagnosis of cardiopulmonary disease in a subject. In other words, the more VOC biomarkers used in the present invention, the higher the sensitivity and specificity of the present invention. Thus, the present invention may include detecting 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 55 or more, 60 or more, 65 or more, 70 or more, 75 or more, 80 or more, 85 or more, 90 or more, 95 or more, or all of the VOC biomarkers disclosed herein (e.g., the biomarkers disclosed in FIG. 16, or the biomarkers specific to each cardiopulmonary disease disclosed herein). Preferably, the present invention includes determining the presence of 5 or more, or 10 or more VOC biomarkers.
[0056] Detecting the presence of a biomarker may include detecting the presence, absence, or level of a biomarker. Detecting the presence of a biomarker may include detecting the level of a biomarker. Detecting the presence or level of a biomarker may include determining the concentration of the biomarker in a sample.
[0057] It will be appreciated that the presence and / or concentration of VOCs can be detected or determined using any suitable method / technique / technology known in the art, such as two-dimensional gas chromatography coupled with mass spectrometry (GC×GC-MS), gas chromatography-ion mobility spectrometry (GC-IMS) technology, gas chromatography (GC), gas chromatography-mass spectrometry (GCMS), mass spectrometry (MS), ion mobility spectrometry (IMS), differential mobility spectrometry (DMS), optical absorption spectrometry, field asymmetric ion mobility spectrometry (FAIMS), electronic nose, selected ion flow tube mass spectrometry (SIFT-MS), protein transfer reaction MS, absorbance / non-dispersive infrared and gas sensors (individually or in an array). Preferably, the detection of the presence and / or concentration of VOCs in a sample of breath from a subject comprises two-dimensional gas chromatography coupled with mass spectrometry (GC×GC-MS). Using GCxGC-MS to detect the presence and / or concentration of VOCs allows for definitive separation of VOC biomarkers, allowing for confident identification of the VOC biomarkers.
[0058] It will be understood that the sample can be analyzed immediately after it is taken from the subject (i.e. it may be a fresh sample). The sample may be placed in a sealed container such as a universal or bijou. The sample may be stored. Preferably, the sample is stored in a sealed / sealable container such as a tube, universal or bijou. Preferably, the container comprises / contains an adsorbent material. Thus, the container may be a sealable container (e.g. a tube) that comprises / contains an adsorbent. The sample can be stored for up to 48 hours. The sample may be stored at a temperature between about 2°C and about 8°C, or between about 3°C and about 6°C. Preferably, the sample is stored at a temperature of about 4°C. Thus, the sample may be stored at a temperature between about 2°C and about 8°C, between about 2°C and about 5°C, between about 3°C and about 6°C, or at a temperature of about 4°C for about 48 hours.
[0059] The sample may be dry purged to reduce the water content of the sample to less than 2 mg per tube. Dry purging may be performed by purging the sample with nitrogen gas. Preferably, dry purging (e.g., dry purging using nitrogen gas) is performed within 48 hours after the sample is collected from the subject.
[0060] "Selecting a subject for treatment" may refer to recording the subject's name and / or identifier so that a third party can recognize that the subject should be treated with a cardiopulmonary therapeutic agent or composition.
[0061] The term "recording" refers to fixing or storing in writing (such as typed) or digitally (such as video or audio recording, or on a computer).
[0062] The subject may be a person suspected of having a cardiopulmonary disease (e.g., asthma, COPD, heart failure and / or pneumonia). Preferably, the subject is experiencing shortness of breath. The term "shortness of breath", also known as dyspnea, refers to difficulty in breathing. This may manifest as rapid shallow breathing, noisy breathing, wheezing, or using shoulder and / or upper chest muscles to help breathe.
[0063] A "subject" may be a vertebrate, a mammal, or a domestic mammal. Thus, the method according to the invention may be used to diagnose or treat any animal, for example, a pig, a cat, a dog, a horse, a sheep, or a cow. Preferably, the subject is a human.
[0064] Some or all of the steps of the methods of the invention may be performed in vitro, ex vivo, or in vivo.
[0065] The method according to the invention may comprise providing a sample obtained from a subject. Thus, the term "exhaled air / breath sample" refers to gas and / or liquid exhaled by a subject, preferably gas and / or liquid (condensate) exhaled from the subject's lungs. The sample is exhaled from the nose and / or mouth of the subject. Preferably, the sample is an exhaled gaseous sample. Thus, the method of the invention does not have to be performed on a subject. The amount of sample may be any amount that provides sufficient biomarker to be measured, for example, the sample may be 500mL to 1L.
[0066] The term "treating" refers to preventing, eradicating, or reducing the severity of cardiopulmonary disease. Thus, the therapeutic agent or composition referred to herein may be any agent that prevents, eradicates, or reduces the severity of asthma, COPD, heart failure, or pneumonia.
[0067] The term "comprising" may refer to "consisting of" or "consisting essentially of."
[0068] All of the embodiments and features described in this specification (including the accompanying claims, abstract and drawings), and / or all steps of any method or process so disclosed, may be combined with any of the above aspects or embodiments, unless otherwise stated with reference to a specific combination, e.g., a combination in which at least some of such features and / or steps are mutually exclusive.
[0069] For a better understanding of the present invention, and to show how embodiments of the same may be practiced, reference is made by way of example to the accompanying drawings in which: [Brief description of the drawings]
[0070] [Figure 1] A visual summary representing the proposed breath testing and diagnostic pipeline. Acute breathlessness patients with cardiorespiratory exacerbations are currently triaged at admission by clinical assessment, digital pathology, and blood biomarkers. Exhaled volatile organic compound biomarkers from the lower airways are visualized using state-of-the-art GCxGC mass spectrometry and undergo a process combining chemical measurements and translational modeling. The resulting exhaled metabolic signatures show specific VOC profiles and VOC classes and colocalization of individual exacerbation subgroups, providing accurate disease classification of acute cardiorespiratory patients.
[0071] [Figure 2-1]Topological data analysis (TDA) representing different acute disease groups annotated by blood biomarkers. Each circle or "node" on the TDA graph represents a subject or group of subjects. Similar subjects are grouped in the same node, the relative similarity of subjects is represented by the node's closeness, and the size of each node is determined by the number of subjects in the node. (A) Visual mapping of acute disease groups in the discovery cohort (n=139) based on 805 discriminatory features, colored by the proportion of acute COPD exacerbations in each node. (B) The network is colored by the mean value of CRP in each node in the discovery cohort (n=139). Higher CRP values were topologically consistent with COPD and pneumonia patients. (C) The network is colored by the mean value of BNP in each node in the discovery cohort (n=139). Higher BNP values were topologically consistent with heart failure patients. (D) The network is colored by the proportion of acute COPD exacerbations in each node in the replication cohort (n=138). In the replication cohort, subjects with pneumonia and COPD exacerbations occupied the extremes of the same TDA network. (E) The network is colored by the mean CRP value at each node. High CRP values were topologically consistent with pneumonia patients. (F) The network is colored by the mean BNP value at each node. High BNP values were topologically consistent with heart failure subjects. [Figure 2-2]Topological data analysis (TDA) representing different acute disease groups annotated by blood biomarkers. Each circle or "node" on the TDA graph represents a subject or group of subjects. Similar subjects are grouped in the same node, the relative similarity of subjects is represented by the node's closeness, and the size of each node is determined by the number of subjects in the node. (A) Visual mapping of acute disease groups in the discovery cohort (n=139) based on 805 discriminatory features, colored by the proportion of acute COPD exacerbations in each node. (B) The network is colored by the mean value of CRP in each node in the discovery cohort (n=139). Higher CRP values were topologically consistent with COPD and pneumonia patients. (C) The network is colored by the mean value of BNP in each node in the discovery cohort (n=139). Higher BNP values were topologically consistent with heart failure patients. (D) The network is colored by the proportion of acute COPD exacerbations in each node in the replication cohort (n=138). In the replication cohort, subjects with pneumonia and COPD exacerbations occupied the extremes of the same TDA network. (E) The network is colored by the mean CRP value at each node. High CRP values were topologically consistent with pneumonia patients. (F) The network is colored by the mean BNP value at each node. High BNP values were topologically consistent with heart failure subjects.
[0072] [Figure 3-1] (A) Scatter plot showing significant differences between exhaled VOC biomarker score values in patients with acute cardiopulmonary disease compared to healthy volunteers. Black horizontal lines in the scatter plot represent the median biomarker score. Mann-Whitney test p-value <0.0001. (B) Receiver operating characteristic (ROC) curves for discovery participants (black line) - AUC1.00 (1.00-1.00), and replication cohort (blue line) - AUC0.89 (0.82-0.95) p<0.0001. (C) Histogram showing the number of patients with high diagnostic uncertainty (blue bar with values above the upper quartile of 20 mm). (D) ROC curve assessing the discriminatory power of exhaled VOC in participants with high diagnostic uncertainty. AUC0.96 (0.92-0.99) p<0.0001. [Figure 3-2] (A) Scatter plot showing significant differences between exhaled VOC biomarker score values in patients with acute cardiopulmonary disease compared to healthy volunteers. Black horizontal lines in the scatter plot represent the median biomarker score. Mann-Whitney test p-value <0.0001. (B) Receiver operating characteristic (ROC) curves for discovery participants (black line) - AUC1.00 (1.00-1.00), and replication cohort (blue line) - AUC0.89 (0.82-0.95) p<0.0001. (C) Histogram showing the number of patients with high diagnostic uncertainty (blue bar with values above the upper quartile of 20 mm). (D) ROC curve assessing the discriminatory power of exhaled VOC in participants with high diagnostic uncertainty. AUC0.96 (0.92-0.99) p<0.0001.
[0073] [Figure 4] (A) Pearson correlations of disease-specific VOC scores and blood-based biomarkers. Pearson correlations demonstrating positive and negative correlations between exhaled VOC scores and blood-based biomarkers. *Significant correlation, p-value <0.05; and (B) Pearson correlations of disease-specific VOC scores and admission observations. Pearson correlations between VOC biomarker scores and admission vital signs. VAS: Visual Analogue Scale (100mm), participants were asked to rate their shortness of breath on a 100mm VAS upon admission.
[0074] [Figure 5-1](A) Circular correlation tree generated based on metabolite set enrichment and chemical similarity analysis of 101 exhaled breath volatiles associated with acute breathlessness. Branches indicate metabolite sets derived using ChemRICH (Methods). Bars show -log10(p) and log2(fold change) values of 101 features extracted using LASSO regression. Figure 16 shows in acute breathlessness compared to control group. Arcs represent Leuven clusters derived from the correlation graph (green if upregulated, red if not significant, and blue if downregulated according to the results of the KS test). Chemical names are color-coded based on chemical classification, and color-coded regions are used to summarize broader chemical groups. (B) Correlation graph showing metabolite communities identified using Leuven clustering. The identity and location of clusters significantly enriched in heart failure are projected onto the circular dendrogram. (C) i) Example GCxGC chromatogram showing the complex profile of exhaled metabolites, ii) 3D rendering of the chromatogram showing visualization of exhaled markers, and iii) phenotypic differences based on features included in the risk score (yellow, asthma; red, pneumonia; magenta, COPD; cyan, heart failure). [Figure 5-2](A) Circular correlation tree generated based on metabolite set enrichment and chemical similarity analysis of 101 exhaled breath volatiles associated with acute breathlessness. Branches indicate metabolite sets derived using ChemRICH (Methods). Bars show -log10(p) and log2(fold change) values of 101 features extracted using LASSO regression. Figure 16 shows in acute breathlessness compared to control group. Arcs represent Leuven clusters derived from the correlation graph (green if upregulated, red if not significant, and blue if downregulated according to the results of the KS test). Chemical names are color-coded based on chemical classification, and color-coded regions are used to summarize broader chemical groups. (B) Correlation graph showing metabolite communities identified using Leuven clustering. The identity and location of clusters significantly enriched in heart failure are projected onto the circular dendrogram. (C) i) Example GCxGC chromatogram showing the complex profile of exhaled metabolites, ii) 3D rendering of the chromatogram showing visualization of exhaled markers, and iii) phenotypic differences based on features included in the risk score (yellow, asthma; red, pneumonia; magenta, COPD; cyan, heart failure).
[0075] [Figure 6] FIG. 13 is a consort diagram outlining acute study recruitment and the number of analyzable GCxGC-MS breath samples.
[0076] [Figure 7] Flowchart showing the removal of breath features from 805 to 101. Due to the high ratio of variables to subjects and potential correlations among the candidate features, Least Absolute Shrinkage and Selection Operator (LASSO) and Elastic Net regularized regression model were adopted as feature selection methods.
[0077] [Figure 8-1]Graphical representation of the probability distribution of the last 101 breath features in the GCxGC-MS peak table. The features mostly follow a similar distribution. Some features contained a mixture of zero and non-zero values, which occurred because measurements were below the lower detection limit of the instrument. Constant features (all zero values) were removed before fitting the main model. [Figure 8-2] Graphical representation of the probability distribution of the last 101 breath features in the GCxGC-MS peak table. The features mostly follow a similar distribution. Some features contained a mixture of zero and non-zero values, which occurred because measurements were below the lower detection limit of the instrument. Constant features (all zero values) were removed before fitting the main model. [Figure 8-3] Graphical representation of the probability distribution of the last 101 breath features in the GCxGC-MS peak table. The features mostly follow a similar distribution. Some features contained a mixture of zero and non-zero values, which occurred because measurements were below the lower detection limit of the instrument. Constant features (all zero values) were removed before fitting the main model. [Figure 8-4] Graphical representation of the probability distribution of the last 101 breath features in the GCxGC-MS peak table. The features mostly follow a similar distribution. Some features contained a mixture of zero and non-zero values, which occurred because measurements were below the lower detection limit of the instrument. Constant features (all zero values) were removed before fitting the main model.
[0078] [Figure 9] 2D visualization of the high-dimensional peak table before adjusting for batch effect. Clustering by collection date "Batch_ID" in the first panel is clearly visible compared to other variables where batch effect is not evident (operator, collection time, wet and dry storage time, and collection volume).
[0079] [Figure 10]2D visualization of the high-dimensional peak table after adjusting for the collection date "Batch_ID". After parametric empirical Bayesian adjustment, no clustering is visible.
[0080] [Figure 11] A) Correlation graph showing how exhaled metabolites (panels of 101) are correlated within each abbreviated subgroup, color-coded based on the Leuvain cluster to highlight differences across networks. Highlighted visual differences include the green Leuvain cluster, which is highly compact in the control group but dispersed in the acute group. B) Output of ChemRICH analysis, showing metabolite sets (circles) significantly enriched during acute breathlessness (size indicating fold change, red = upregulated, blue = downregulated). The upregulated metabolite set with high chemical similarity (based on Tanimoto coefficient) consisted mainly of acyclic and branched hydrocarbons and belonged to the green Leuvain cluster (indicated by the color of the outer ring). The quantitative output of the ChemRICH analysis complements the visual differences in the graph network.
[0081] [Figure 12-1] Violin plots showing significant differences between VOC biomarker score values across different disease subgroups are shown. *Kruskal-Wallis test for comparing non-parametric data. *Significant p-value <0.0001. [Figure 12-2] Violin plots showing significant differences between VOC biomarker score values across different disease subgroups are shown. *Kruskal-Wallis test for comparing non-parametric data. *Significant p-value <0.0001. [Figure 12-3] Violin plots showing significant differences between VOC biomarker score values across different disease subgroups are shown. *Kruskal-Wallis test for comparing non-parametric data. *Significant p-value <0.0001. [Figure 12-4]Violin plots showing significant differences between VOC biomarker score values across different disease subgroups are shown. *Kruskal-Wallis test for comparing non-parametric data. *Significant p-value <0.0001. [Figure 12-5] Violin plots showing significant differences between VOC biomarker score values across different disease subgroups are shown. *Kruskal-Wallis test for comparing non-parametric data. *Significant p-value <0.0001.
[0082] [Figure 13-1] Kaplan-Meier survival analysis is shown. (A) A total of 29 patients were readmitted within 60 days after discharge. (B) Total number of readmitted patients categorized by median VOC score for acute illness showing no significant difference in readmission rates based on underlying VOC score p-value 0.77 (log-rank test for equality of survival functions). (C) Total number of deaths during the 2-year follow-up period (n=12). (D) Kaplan-Meier survival analysis for all-cause 2-year mortality categorized by disease group. (E) Kaplan-Meier survival analysis for all-cause 2-year mortality categorized by median VOC score for acute illness. [Figure 13-2] Kaplan-Meier survival analysis is shown. (A) A total of 29 patients were readmitted within 60 days after discharge. (B) Total number of readmitted patients categorized by median VOC score for acute illness showing no significant difference in readmission rates based on underlying VOC score p-value 0.77 (log-rank test for equality of survival functions). (C) Total number of deaths during the 2-year follow-up period (n=12). (D) Kaplan-Meier survival analysis for all-cause 2-year mortality categorized by disease group. (E) Kaplan-Meier survival analysis for all-cause 2-year mortality categorized by median VOC score for acute illness.
[0083] No significant difference between groups, p-value 0.07 (log-rank test for equality of survival functions).
[0084] [Figure 14] FIG. 13 is a graph showing the overall classification accuracy using all five biomarker scores.
[0085] [Figure 15] (A) Comparative ROC analysis demonstrating the diagnostic value of the asthma VOC score for acute disease subgroups primarily due to infection (pneumonia and COPD) in the pooled (discovery and replication) cohort. (B) Comparative ROC analysis demonstrating the diagnostic value of the heart failure VOC score for other acute disease subgroups (asthma, COPD, and pneumonia) in the pooled cohort.
[0086] [Figure 16-1] Chemical names, CAS Registry Numbers, KEGG, Human Metabolome Database and ChEBI identifiers, and MSI-based metabolite identification levels, concentration ranges, and fold changes (expressed as log2) in acute and control groups, as well as chemical assignments of selected predictions from regression models detailing the contribution of compounds to disease-specific biomarker risk scores (†adjusted p-values <0.05). [Figure 16-2] Chemical names, CAS Registry Numbers, KEGG, Human Metabolome Database and ChEBI identifiers, and MSI-based metabolite identification levels, concentration ranges, and fold changes (expressed as log2) in acute and control groups, as well as chemical assignments of selected predictions from regression models detailing the contribution of compounds to disease-specific biomarker risk scores (†adjusted p-values <0.05). [Figure 16-3] Chemical names, CAS Registry Numbers, KEGG, Human Metabolome Database and ChEBI identifiers, and MSI-based metabolite identification levels, concentration ranges, and fold changes (expressed as log2) in acute and control groups, as well as chemical assignments of selected predictions from regression models detailing the contribution of compounds to disease-specific biomarker risk scores (†adjusted p-values <0.05). [Figure 16-4]Chemical names, CAS Registry Numbers, KEGG, Human Metabolome Database and ChEBI identifiers, and MSI-based metabolite identification levels, concentration ranges, and fold changes (expressed as log2) in acute and control groups, as well as chemical assignments of selected predictions from regression models detailing the contribution of compounds to disease-specific biomarker risk scores (†adjusted p-values <0.05). [Figure 16-5] Chemical names, CAS Registry Numbers, KEGG, Human Metabolome Database and ChEBI identifiers, and MSI-based metabolite identification levels, concentration ranges, and fold changes (expressed as log2) in acute and control groups, as well as chemical assignments of selected predictions from regression models detailing the contribution of compounds to disease-specific biomarker risk scores (†adjusted p-values <0.05). [Figure 16-6] Chemical names, CAS Registry Numbers, KEGG, Human Metabolome Database and ChEBI identifiers, and MSI-based metabolite identification levels, concentration ranges, and fold changes (expressed as log2) in acute and control groups, as well as chemical assignments of selected predictions from regression models detailing the contribution of compounds to disease-specific biomarker risk scores (†adjusted p-values <0.05).
[0087] [Figure 17] FIG. 11 is a Venn diagram showing the distribution of the final panel of 101 breath 362 breath biomarkers across different disease groups. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0088] Working Example Disclosed herein is a real-world, prospective study of acutely unwell hospitalized patients with shortness of breath due to severe exacerbations of cardiopulmonary disease (asthma, chronic obstructive pulmonary disease (COPD), heart failure or pneumonia) and healthy controls. It is demonstrated that respiratory biomarkers can reliably and repeatedly identify acute cardiopulmonary shortness of breath; this includes cases where diagnostic uncertainty exists.
[0089] method Study design, participants and ethical approval This clinical study was a prospective, real-world, observational study conducted at a tertiary cardiorespiratory centre in Leicester, UK. Participants were recruited year-round from May 2017 to December 2018.
[0090] Patients presenting to University Hospital Leicester (UHL) with self-reported acute breathlessness and requiring hospital admission and / or change in baseline treatment were invited to participate in the study. After triage and advanced clinical assessment, if (i) acute decompensation of heart failure, (ii) exacerbation of asthma / COPD, or (iii) an initial clinical diagnosis of adult community-acquired pneumonia was suspected by the triage nurse / attending clinician at the time of triage, a member of the study team would assess the patient against predefined eligibility criteria for study participation. Informed consent was obtained from all participants within 24 hours after admission. Healthy volunteers matched for age and / or home environment were recruited. Where environmentally matched controls were inappropriate, healthy volunteers were recruited through local recruitment databases and advertisements. Details of comorbidities and medication use of the healthy volunteers are shown in Table 1.
[0091] This trial was conducted in accordance with the ethics and principles of the Declaration of Helsinki and Good Clinical Practice guidelines. All patients provided written consent. The National Research Ethics Service Committee East Midlands approved the study protocol (REC number: 16 / LO / 1747). Integrated Research Approval Scheme (IRAS) 198921.
[0092] [Table 1-1] [Table 1-2]
[0093] Recruitment began in February 2017, followed by analytical method development and optimization of a robust sample pathway to enable ongoing deployment, collection and analysis of sorbent tubes to and from the clinic, and sample analysis by GCxGC-MS was set up and brought online later that year (August 2017). While the overall study denominator was 455 participants, the GCxGC-MS study presented here had 363 participants, resulting in a GCxGC-MS completion rate of 76% (Figure 6).
[0094] A detailed survival analysis of the study participants is shown in Figure 13.
[0095] This trial was conducted in accordance with the ethics and principles of the Declaration of Helsinki (deceleration) and Good Clinical Practice guidelines. All patients provided written consent. The National Research Ethics Service Committee East Midlands approved the study protocol (REC number: 16 / LO / 1747). Integrated Research Approval Scheme (IRAS) 198921.
[0096] clinical judgment A clinical adjudication process was implemented to precisely define and quantify diagnostic labels in the study and address potential misclassification. A panel of two senior clinical adjudicators (SS and NG) reviewed all available case records and images and determined the primary diagnosis for each case by discussion to reach consensus. The degree of diagnostic uncertainty was marked on a 100 mm visual analogue scale (VAS scale), blinded to the given diagnosis and blood biomarkers.
[0097] This process was implemented with a focus on mirroring the acute triage pathway for all pathology data required to support the diagnosis (e.g., CRP, BNP not available for initial clinical review).
[0098] The degree of diagnostic uncertainty obtained from the clinical adjudication process was taken into account for block randomization, and subjects with higher diagnostic uncertainty (≥ upper quartile = 20 mm) were assessed separately as previously described ( Fig. 3c–d ).
[0099] Collection of breath samples Breath sampling was attempted for all consenting participants using the CE marked breath sampling device "Breath Collector for Extracorporeal Analysis" RECIVA® (Owlstone Nanotech Ltd) in combination with a dedicated clean air supply unit. The ReCIVA® device aims to standardize the collection of alveolar breath by providing patients with a clean air supply, devoid of VOCs. The flow rate, volume and proportion of breath collected are controlled while exhaled VOCs are sampled directly into the sorbent tube. The ReCIVA® setting mode was set to "lower airways only" and continuous monitoring of CO2 and partial pressures allowed targeting the alveolar portion, which is rich in exhaled VOCs. The collection volume, flow rate and maximum sampling time were 1 L, 250 mL min , respectively. -1 , and 900 seconds. Breath sampling was well tolerated by all participants.
[0100] During sampling, room air and supply air were also sampled as environmental controls. This involved attaching the sorbent tube to a handheld personal pump (Escort Elf, Sigma Aldrich, Dorset, UK) and either leaving the sampling end open to room air or connecting it via a T-piece to a ReCIVA® air supply line. 0.5 L min -1 A total of 1 L of air was collected for 2 minutes at a flow rate of 0.01 µm.
[0101] Sorbent tubes were immediately capped (brass cap, Markes International Ltd), placed in a refrigerator at 4°C, and shipped to the laboratory within 72 hours. To minimize background variation, sample collection was completed in the same procedure room attached to the inpatient unit where possible, but for unwell patients or those requiring supplemental oxygen, samples were taken at the bedside.
[0102] Sample storage and preparation On arrival, samples were purified with nitrogen (CP grade with in-line trap, BOC, Leicester, UK) at a flow rate of 50 mL min -1 The samples were then dried and purged for 2 min using a 20 μg mL aliquot of deuterated toluene and octane. -1 A 0.6 μL aliquot of the standard solution was poured into a tube at a flow rate of 100 mL min -1 The mixture was then added to a stream of nitrogen for 2 minutes to purge excess solvent.
[0103] Analysis of indoor and supply air samples Two separate elastic net regression models were fitted to the peak tables for the indoor and supply air samples. [ka] was transformed and adjusted for batch effects (date of collection) using PEBA. The independent variables were the final set of 101 features, and the dependent variable was the clinical diagnosis (acute asthma, acute COPD, pneumonia, heart failure, or healthy volunteers). After 100 iterations of 10-fold cross-validation for each of the two models, only two features were found to have stable non-zero regression coefficients. These features concern supply air, a component of the pneumonia score, and indoor air, a component of the health score, highlighting the robustness of the selected feature separation model.
[0104] Breath analysis TD-GC×GC-FID / MS Breath samples were analyzed by thermal desorption using flow modulation and comprehensive two-dimensional gas chromatography (GCxGC) coupled with dual flame ionization detection and mass spectrometry (MS). Dual detection using MS and flame ionization detection (FID) provides both quantitative and qualitative results, utilizing excess flow from a flow-based modulator suitable for the analysis of volatiles.
[0105] Analysis by GC×GC was optimized and performed using an Agilent 7890A gas chromatogram equipped with a 5799B mass spectrometer equipped with a CFT flow modulator and a high-efficiency EI ion source (Agilent Technologies Ltd, Stockport, UK). The instrument was interfaced with a TD-100xr thermal desorption autosampler (Markes International Ltd, Llantrisant, UK). Samples were analyzed in trays; typically six per tray plus a reference mixture containing n-alkanes and aromatics run per tray, and a reference indoor air VOC mixture run every fourth tray. Data were acquired with MassHunter GC-MS Acquisition B.07.04.2260 (Agilent) and processed (i.e. baseline correction, alignment, feature extraction) with a pre-developed and optimized workflow using GC Image® v2.8 suite (GC Image, LLC. Lincoln, NE, US) and Python. The sorbent tubes used were Tenax / TA equipped with a Carbograph 1TD (Hydrophobic, Markes International Ltd) equipped with a corresponding cold trap. Chromatographic features resulting from analytical artefacts were removed from the peak table (e.g. ubiquitous siloxanes).
[0106] For quality control purposes, samples were analyzed using a detailed sample history, and metadata and experimental data were recorded at all stages of collection and analysis using the open-access LabPipe toolkit.
[0107] Chemical speciation of identified breath biomarkers The chemistry of volatile metabolites exhaled in breath consists of a diverse mixture of non-novel low molecular weight compounds. Thus, for the majority of features, chemical identification involved comparison to authentic reference compounds according to the Metabolomics Standards Initiative (MSI) Level 1 criteria for metabolite identification (Figure 16). Identification was based on at least two independent orthogonal identifiers, including primary and secondary retention times, mass spectral similarity matches, and calculated retention indices. Where authentic reference compounds were not available, chemical identification was in accordance with MSI Level 2 for putative annotation. The highly structured chromatographic data and group type separation provided by GCxGC, alongside a well-characterized chromatographic space from the analysis of an extensive library of authentic compounds, increased confidence in the tentative assignments. The GCxGC orthogonal separation also meant that chemical identification of unknown metabolites could be made in accordance with at least MSI Level 3 for putative chemical classification.
[0108] The diagnostic accuracy of reported breath VOCs was tested according to the reporting standards of the Diagnostic Accuracy Study Guidelines, and for multivariate prediction models, the Transparency Reporting of Multivariate Prediction Models for Individual Prognosis or Diagnosis (TRIPOD) was followed.
[0109] Quality control and quality assurance system A number of traceable and verifiable quality control and quality assurance (QC / QA) procedures are applied throughout the breath sampling and analysis process, which effectively prevents potential defects and ensures high standards of deliverables.
[0110] Four criteria were used to select high quality breath samples in order to eliminate poor quality samples from the final analysis. These were: 1. Collect at least 800 mL of breath from the patient to adequately pre-enrich trace VOCs present in the breath. 2. Concentrations of isoprene and acetone in the supply air were less than three standard deviations from the mean supply air concentration, which prevented the incorrect assignment of breath samples as supply air samples. 3. The concentrations of isoprene and acetone in the exhaled breath were more than 10 and 5 standard deviations, respectively, above the levels measured in the patient's air supply, ensuring that the samples were collected in the sorbent tube and not an incorrectly assigned air supply sample. 4. The chromatogram was visually inspected and found to be not distorted by large amounts of exogenous compounds (i.e., overloaded peaks).
[0111] An overview of the number of breath samples fulfilling all QC / QA criteria is shown in (Figure 6).
[0112] Sample Analysis QC / QA Procedures For quality control purposes, samples were analyzed following a previously published workflow, and detailed sample history, metadata, and experimental data were recorded at all stages of collection and analysis using the open-access LabPipe toolkit. Chromatographic methods were optimized with respect to peak shape, sensitivity, and separation. The stability of the TD-GCxGC-FID / MS analysis was tracked using quality control charts of internal standards, and instrument performance was evaluated following evaluation of the variation in retention time, peak area, and shape of VOCs in two standard reference mixtures. For every 6 samples. The number of thermal cycles and weight of each tube were recorded to monitor tube age and integrity before conditioning and sending to the clinic. After every conditioning cycle, all tubes were given a batch number, and batch blanks were analyzed to monitor contamination from the beginning of the sample preparation process. In addition, all batches were given a 2-week expiration date to ensure regular monitoring. To minimize the influence of biological and analytical confounding factors (e.g., circadian rhythm, sample stability), potential influences by the operator, analysis date, collection time, storage time before dry purge, sample storage time after dry purge, and collection volume were evaluated and taken into account in batch modifications if necessary. In addition to the routine analysis of reference standards used to monitor retention shifts and instrument response, the TD-GCxGC analytical system ran programmed thermal cycles between each sample to mitigate potential problems arising from sample carryover, and TD trap blanks and empty sorbent tubes were analyzed every sixth sample to monitor the instrument baseline signal.
[0113] Statistical procedures Statistical analyses were performed using R (3.6.1 and 4.0.0, R Core Team (2019)). This study used the SPECTRE high performance computing facility at the University of Leicester. Baseline data and values were expressed as mean ± (SD) and median (IQ range). Data were analyzed using ANOVA to assess differences between groups for normally or approximately normally distributed variables and Kruskal-Wallis for non-normally distributed variables. Pearson's chi-square and Fisher's exact were used to assess differences in categorical variables. All P values are two-sided and significant at the 0.05 level unless otherwise reported. Study sample size calculations were performed based on sample size estimates of appropriate sensitivity and / or specificity (sample size estimation section).
[0114] Discovery and Replication Sets Two hundred seventy-seven subjects were post-hoc randomized into discovery and replication cohorts in a 1:1 ratio through block random allocation. Randomization was stratified based on (I) determined clinical diagnosis, (II) time from admission to breath test, and (III) clinical diagnostic uncertainty score. Block random allocation was performed using the R package randomizer. After block randomization, the discovery and replication sets contained 139 and 138 subjects, respectively.
[0115] Checking topological equivalence in discovery and replication sets Topological data analysis is an unsupervised machine learning tool used to analyze large, high-dimensional, complex datasets. It is highly sensitive to patterns that are often overlooked by other data reduction tools such as principal component analysis (PCA).
[0116] TDA captures the shape of the data, providing a meaningful geometric representation in which complex relationships among data points can be preserved and considered together.
[0117] Before performing TDA, each feature [ka] transformed. TDA parameters were set as follows: number of hypercubes = 20, where the number of hypercubes refers to the number of overlapping intervals in the projection. Distance between data points was measured using Euclidean distance. The first two linear discriminant functions (LD1) and (LD2) were used as projections. Clustering of overlapping intervals on the projections was performed using agglomerative (bottom-up) hierarchical clustering with complete linkage. TDA was performed using Kepler Mapper 1.4.0 and Python 3.5.
[0118] Here, we calculated equivalence between topological data shapes generated using 805 volatile features extracted from GCxGC-MS peak tables, both in discovery and replication cohorts (Figure 2).
[0119] Selection of breath characteristics Feature selection was implemented by Lasso and Elastic-Net Regularized Generalized Linear Model (GLMNET) using the glmnet package in R. [ka] After removing features present in <80% of all samples from the transformed discovery GCxGC-MS peak table, a matrix of 735 features was obtained. A multinomial regression model with LASSO regularization was fitted to the 735 feature matrix of the discovery set using 10-fold cross-validation with clinical diagnosis (acute asthma, acute COPD, pneumonia, heart failure, or healthy volunteer) as the dependent variable of the model. The 10-fold cross-validation was repeated 100 times, and features with non-zero regression coefficients in more than 80 cross-validation runs were considered as stable candidate features predicting the outcome (clinical diagnosis), resulting in 278 stable candidate features.
[0120] Multinomial regression models with elastic net regularization were fitted to the 278 features with clinical diagnosis as the dependent variable in the model. After chemometric testing and lasso and elastic regression analyses detailed above, a final set of 101 breath volatile compounds was generated (Figure 7).
[0121] A multinomial regression model with elastic net regularization was fitted to the matrix of 101 exhaled biomarkers with 100 iterations of 10-fold cross-validation. The R package glmnetUtils was used to determine the optimal value of α, the elastic net penalty. The optimal value of α was 0 (ridge regression). A linear combination of the most stable features from the multinomial regression models fitted to the 101 biomarkers formed a set of scores for predicting the probability of belonging to different disease groups (acute asthma, acute COPD, pneumonia, heart failure, or healthy volunteers). A ridge regression with a logit link function (binary logistic regression) was fitted to the 101 respiratory-related features, with the dependent variable being “acute illness” as the binary outcome. The linear predictor resulting from the combination of the most stable features was used as the score for predicting acute illness.
[0122] Coexpression and functional enrichment analysis It was of interest to investigate whether there existed sets of "co-expressed" features within the final set of 101 features, i.e. sets that contain correlated features. Considering sets of co-expressed features has value in terms of reducing the dimensionality of the problem and mitigating the problem of multiple testing using enrichment scores. Co-expression and functional enrichment analyses are described in (Supplementary Information).
[0123] Metabolite sets were derived based on Ward hierarchical cluster analysis using the ChemRICH method (Figure 5A), while the broader community was derived from a Louvain cluster analysis to aid in the interpretation of correlation graphs (Figure 5B, see Supplementary Information section on co-expression and functional enrichment analysis). As covariation between metabolites has no evidential value in itself, set-level significance was established using the Kolmogorov-Smirnov test (KS test) using the ChemRICH method, and the Tanimoto coefficient was calculated to assess chemical similarity within the set using Metabox, with frequency of occurrence in published literature and relevant databases taken into account (KEGG, ChEBI, Human Metabolome Database, Human Breath Omics Database and Microbial VOC Database). Chemical similarities are of interest because compounds derived from similar pathways may share common structural features or chemical groups. This combined data-driven and chemistry-driven approach is shown to improve enrichment analysis and enable further interpretation of the central findings herein (Figure 11).
[0124] Supplementary Information (SI) Probability distribution of respiratory features (biomarkers): Features in the GCxGC peak table fell into three broad categories: (1) constant features (values of zero for all samples), (2) features that contained a mixture of zero and nonzero values, and (3) features that contained all nonzero values. Zero values occurred because measurements were below the lower detection limit of the instrument. Constant features were removed prior to fitting the main model.
[0125] The graphical distribution of the final 101 features (biomarkers), which mainly fall into type 2 and type 3 categories, is shown in (Figure 8). A spike in 0 values is clearly visible for certain features. Based on these observations, a reasonable choice as a theoretical model for the probability distribution of features from a GCxGC-MS peak table could be the zero-corrected log-normal distribution.
[0126] Mitigating the adverse impact of batch effects in biomarker pattern detection Batch effects are a common problem in omics data analysis. The presence of batch effects makes it difficult to compare data collected and analyzed at different processing times (Figures 9 and 10).
[0127] The following factors were investigated as possible sources of batch variation: I.Batch_ID - Date of sample collection: (1) Batch 1 – August 2017 to October 2017 (2) Batch 2 – November 2017 to March 2018 (3) Batch 3 – April 2018 to December 2018 II. Operators: (N: 1-6) - Refers to the members of the research team who will operate RECIVA throughout the entire sampling program. III. Time of day when sample was collected (circadian rhythm): (1)1 = 9:00 a.m. to 11:00 a.m. (2)2 = 11:00 a.m. to 1:00 p.m. (3) 3 = 1 pm to 3 pm (4) 4 = 3:00 p.m. to 5:00 p.m. IV. Thyme samples kept moist (1) 1 = 0 to 2 days (2) 2 = 2 to 5 days (3) 3 = 5 to 10 days (4) 4 = 10 to 20 days (5) 5 = 20 to 42 days (6) 6 = 42 days or more V. Time stored in dry conditions (after dry purging) (1) 1 = 0 to 2 days (2) 2 = 2 to 5 days (3) 3 = 5 to 10 days (4) 4 = 10 to 20 days (5) 5 = 20 to 42 days (6) 6 = 42 days or more VI. Volume of exhaled air collected (above 80% threshold): (1) 1=100% (2) 2=90~99% (3) 3 = 80~89%
[0128] Figure 9 visualizes the GCxGC-MS peak table with all 805 features using t-stochastic nearest neighbor embedding (tSNE). Clustering by "collection date" is visible (top left plot). For the remaining factors, no obvious clustering seems to exist. The effect collection date was adjusted by applying Parametric Empirical Bayesian Adjustment (PEBA). PEBA was performed using the ComBat function from Bioconductor's SVA package. The results of this adjustment are shown in (Figure 10). It can be seen that the clustering by collection date is no longer evident. The peak table adjusted for batch effect was used in all subsequent feature selection models.
[0129] Model Accuracy The overall classification accuracy of the statistical model using all five biomarker scores from the final set of 101 breath features was assessed by comparing the equilibrium accuracy of a model trained using the true class labels to the equilibrium accuracy of the same model tested using randomly shuffled class labels. This process was repeated 1000 times. The overall classification accuracy using all five biomarker scores was 0.722, 95% CI (0.6653-0.774), and the results are shown in Figure 14.
[0130] Chemical speciation of identified breath biomarkers To confirm the chemical identity of the concatenated list of 101 breath peaks, available standard reference compounds were purchased and analyzed, including C8-C20 saturated alkane certified standards (Sigma Aldrich, Dorset, UK), aromatic calibration standards (NJDEP EPH 10 / 08 Rev.2, Thames Restek, Saunderton, UK), a multicomponent room air standard (Sigma Aldrich, Dorset, UK), two terpene reference mixtures (Spex Centriprep, Emerald Scientific, San Luis Obispo, US), and individual standards from Sigma Aldrich (Merck Life Sciences), Greyhound Chromatography, Scientific Lab Supplies, Alfa Chemicals, and SantaCruz Biotechnology.
[0131] Figure 16 lists the chemical assignments of predictive markers selected from the regression models detailing chemical names, CAS Registry Numbers, KEGG, Human Metabolome Database, and ChEBI identifiers, as well as MSI-based metabolite identification levels, concentration ranges, and fold changes (expressed as log2) between acute and control groups, and the contribution of compounds to disease-specific biomarker risk scores (†adjusted p-values < 0.05).
[0132] Sample size estimation The study protocol aimed to recruit 550 subjects, which would allow for the identification of sensitive biomarkers (≥80%) for acute breathlessness with a maximum allowable error in the estimate of sensitivity not exceeding 5% with 95% confidence.Similarly, a specific biomarker for acute breathlessness (≥80%) may be identified with a maximum allowable error in the estimate of specificity not exceeding 5% with 80% confidence, reaching a total sample size of n=277.
[0133] Based on a total sample size of n=277, a posterior sample size calculation was performed using sensitivities of 70% and 80% ± (precision 10%, 15%, and 20%) to obtain biomarkers that could "rule out" the acute disease class. The same target was also applied to specificity. Calculations were performed using a confidence level of 95%.
[0134] Recruitment assumed an 80% prevalence of acute illness, with 1:5 of recruited patients being healthy controls without shortness of breath (Table 2). Although it is acknowledged that the assumption of an 80% prevalence of acute illness places limitations on the validity of sample size calculations, an estimate of an 80% prevalence is not unreasonable based on clinical expectations.
[0135] [Table 2]
[0136] Coexpression and functional enrichment analysis Graph Building and Cluster Analysis Subjects from both the discovery and replication sets, excluding healthy subjects, were analyzed using a data matrix containing 101 features obtained from the previous regression analysis. [ka] The Spearman rank correlation matrix is [ka] was calculated for.
[0137] A scale-free graph g is defined as an adjacency matrix [ka] It was constructed by generating Where: [ka] teeth [ka] is the sample correlation matrix of β, where β≧1.
[0138] We estimated β using the pickSoftThreshold function in the WGCNA package in R. The igrap package in R is [ka] was used to construct g, which is a weighted unsigned graph. We call this graph g the "association graph".
[0139] Then, Louvain clustering was performed on the correlation graph to obtain a set of eight features.
[0140] Eight feature sets obtained from the Louvain clustering on the correlation graph were used for the enrichment analysis. Instead of considering individual features and how they distinguish different disease groups, a set of features was considered, with the idea that combining features may yield better discriminatory power. The enrichment analysis was performed using the Bioconductor (version 3.12) packages GSVA and limma. Feature set 3 was found to be enriched in asthma and HF, while feature set 5 was found to be enriched only in HF; see Tables 3-6. Enriched feature sets 3 and 5 did not demonstrate an improvement in diagnostic accuracy over the scores obtained from the regression analysis.
[0141] [Table 3]
[0142] [Table 4]
[0143] [Table 5]
[0144] [Table 6] EXAMPLES
[0145] Example 1 - Overview We sampled and analyzed exhaled breath from 277 participants recruited from acutely breathless hospitalized patients and matched healthy controls to identify metabolic class dysregulations in cardiopulmonary disease and to investigate whether exhaled VOC profiles can predict acute cardiopulmonary exacerbations despite diagnostic uncertainty, and therefore play a potential role in determining the phenotype of acute cardiorespiratory breathlessness.
[0146] The mean (SD) age of participants was 60.8 ± (16.8) years, 51% were male, 30 patients required supplemental oxygen on admission, and the mean admission modified early warning score (mEWS-2 score) was 2. The cohort consisted of patients presenting with the following exacerbation subtypes; acute severe asthma (n = 65), acute severe COPD (n = 58), acute severe heart failure (n = 44), community-acquired pneumonia (n = 55), and healthy volunteers (n = 55), recruited between May 2017 and December 2018 (Figure 6). The demographic and clinical characteristics of participants are summarized in Table 7. Exhaled breath samples were collected using a ReCIVA® device and analyzed using thermal desorption (TD) coupled with comprehensive two-dimensional gas chromatography, employing a standardized sampling and gating protocol to enrich for alveolar volatiles. (GCxGC) dual flame ionization detection (FID) and mass spectrometry (MS) are used (Figure 1 and Methods).
[0147] [Table 7-1] [Table 7-2] [Table 7-3] EXAMPLES
[0148] Example 2 - Unbiased discovery using topological data analysis identifies respiratory markers of acute illness To achieve unbiased discovery of exhaled breath VOCs predictive of acute disease groups, patients were post hoc randomized and blocked into a discovery cohort of 139 participants (acute asthma n=33, acute COPD n=29, acute heart failure n=22, community-acquired pneumonia n=28, healthy volunteers n=27), and a replication cohort of 138 participants (acute asthma n=32, acute COPD n=29, acute heart failure n=22, community-acquired pneumonia n=27, healthy volunteers n=28). Randomization allowed for internal replication of diagnostic exhaled breath biomarkers while adjusting for relevant confounders. Details of the randomization and further clinical characteristics of the cohorts are described in Methods and Table 1 and Table 8. Chemical analysis and quantification of VOCs were performed blinded to clinical diagnosis by two analytical chemists (MW and RC) and data locking by an independent statistician (MR) was followed by biostatistical analysis linking subject identifiers with chemometric biomarkers.
[0149] Using TD-GCxGC-FID / MS, 805 unique chromatographic features (peaks) were detected across the breath sample set. Applying topological data analysis (TDA) to these 805 chromatographic features yielded topologically distinct networks that differentiated the underlying causes of acute breathlessness while identifying corresponding blood-based biomarkers in both the discovery and replication cohorts (Figure 2). Specifically, healthy volunteers and patients with acute heart failure formed distinct topological groups in both the discovery and replication populations, whereas acute asthma, acute COPD, and respiratory hospitalization due to pneumonia formed a topological continuum, albeit within different regions of a single network in the replication cohort. Similar findings were observed in the discovery cohort, although acute asthma formed a separate group.
[0150] [Table 8-1] [Table 8-2] EXAMPLES
[0151] Example 3 - Biomarker profiling and risk scores To generate a concatenated list of breath biomarkers suitable for diagnostic use, we applied a threshold of 80% feature presence per patient group, below which features were removed (Figure 7). This approach was further supported by the unique distribution characteristics of the breath biomarkers (Figure 8), allowing the generation of patient-specific multi-VOC biomarker risk scores. A further filtering step using the Least Absolute Shrinkage and Selection Operator (LASSO) and ElasticNet regression methods, followed by the removal of 38 peaks considered to be chemical and material artifacts (e.g., siloxanes), generated a final panel of 101 breath volatiles (Figure 7). This analysis plan therefore enabled the identification of abundant and chemically diverse responses in the VOC profile, as opposed to a small number of individual VOC markers, allowing the generation of a biomarker risk score. Data were inspected for batch effects and adjusted accordingly. A batch effect was detected in relation to a major equipment maintenance event (occurred twice creating three groups; see Supplementary Information section on batch adjustment). No significant contribution was observed based on the ReCIVA device used, operator, time of day, or volume of exhaled breath samples collected, most likely nullified by simultaneous and sequential recruitment across all cohorts throughout the study to reduce potential bias (Figures 9-10). Values of the generated VOC biomarker risk score were found to be significantly higher in patients with acute cardiopulmonary disease compared to healthy volunteers (Figure 3a). In the discovery cohort (n=139), the VOC biomarker risk score was able to effectively distinguish participants exhibiting acute cardiopulmonary exacerbation from age-matched healthy controls with an area under the curve (AUC) of 1.00 (1.00-1.00) p<0.0001, sensitivity of 1.00 (1.00-1.00), specificity of 1.00 (1.00-1.00), positive predictive value (PPV) of 1.00 (1.00-1.00), and negative predictive value (NPV) of 1.00-1.00.In the replication cohort (n=138), the same VOC biomarker risk score distinguished acutely ill participants from healthy controls with AUC 0.89 (0.82-0.95) p<0.0001, sensitivity 0.79 (0.71-0.86), specificity AUC 0.85 (0.72-0.98), PPV 0.95 (0.91-0.99), and NPV 0.51 (0.36-0.65) (Figure 3b).
[0152] After the clinical adjudication process (Methods), a degree of clinical diagnostic uncertainty was assigned to each patient using a 100 mm visual analog scale (VAS) at the time of clinical triage (Fig. 3c). Diagnostic uncertainty was defined as patients with VAS values in the top quartile ≥20 mm. The VOC biomarker risk score for acute disease was able to identify acute disease with an AUC of 0.96 (0.92-0.99) p < 0.0001, sensitivity of 0.90 (0.82-0.97), specificity of 0.92 (0.85-0.99), PPV of 0.93 (0.86-0.99), and NPV of 0.89 (0.81-0.97) (Fig. 3d).
[0153] Further comparative receiver operating characteristic (ROC) analysis was performed to evaluate the diagnostic accuracy of the asthma biomarker score for primarily infectious respiratory diseases (pneumonia and COPD) in the pooled cohort curve AUC: 0.70 (0.62-0.78) p<0.0001, sensitivity 0.72 (0.64-0.83), specificity 0.64 (0.55-0.73), PPV 0.54 (0.43-0.64), NPV 0.80 (0.72-0.88). ROC analysis was performed to evaluate the diagnostic value of the heart failure biomarker score relative to other acute disease groups AUC: 0.78 (0.70-0.86) p<0.0001, sensitivity 0.77 (0.64-0.89), specificity 0.71 (0.64-0.78), PPV 0.40 (0.29-0.50), NPV 0.92 (0.88-0.97) (Figure 15). EXAMPLES
[0154] Example 4 - Correlation of Exhaled Biomarker Scores with Blood-Based Biomarkers and Hospitalization Observations As previously described, VOC biomarker risk scores were generated for each of the acute disease subgroups and for healthy subjects without cardiopulmonary breathlessness. In the combined discovery and replication cohorts (n=277), in addition to a significant negative correlation between VOC scores of health status and CRP and BNP (n=277, r=-0.15, p<0.0001 and -0.21, p<0.0001, respectively), there were weak but statistically significant positive correlations between VOC scores of pneumonia and CRP (n=277, r=0.33, p<0.0001) and acute heart failure and BNP (n=277, r=0.33, p<0.0001) (Figure 4a).
[0155] Interestingly, a significant correlation was also identified between the acute illness VOC score and vital signs observed during triage (Figure 4b). EXAMPLES
[0156] Example 5 - Chemical classification of predictive markers in disease groups Chemical identification of the 101 biomarker panel involved comparison with authentic reference compounds that conformed to the Metabolomics Standards Initiative (MSI) Level 1 criteria for metabolite identification (Figure 16).
[0157] The most common chemicals associated with acute shortness of breath in this study were linear and methyl-branched hydrocarbons (30%), ketones (10%), aldehydes (8%), and terpenes (13%), followed by other less prevalent and relevant classes such as sulfur-containing VOCs (7%), alcohols (6%), aromatics (5%), esters (3%), nitrogen-containing VOCs (3%), ethers (2%), halogenated compounds (1%), and various acrylates (12%) (Figure 16). EXAMPLES
[0158] Example 6 - Enrichment of Metabolite Sets and Chemical Similarity Analysis Unlike functional metrics that rely on mapping metabolites with known, well-annotated metabolic pathways, metabolic changes indicative of a response can be derived independently. To obtain clues that may signal responsiveness, we assessed a panel of 101 features across covariation clusters, or metabolite sets (Figure 5A and Figure S11).
[0159] Overall, we identified 20 metabolite sets, 11 of which were enriched during acute cardiorespiratory disease exacerbations. The seven upregulated metabolite sets were composed primarily of acyclic and branched hydrocarbons (sets 3, 5, 7, and 9 in Figure 11). The results of our analysis herein show co-expression of highly chemically similar and significantly enriched hydrocarbons, providing primary evidence of exhaled breath VOCs indicative of disease response measured in vivo. This is clearly illustrated in (Figure 5a), where metabolite sets (inner tree) are labeled by broader chemical classifications (outer rings). C 5-7 , C 8-10 , and C 11-16 cluster based on carbon number, and also shows the greatest change during acute exacerbations. EXAMPLES
[0160] Example 7 - Diagnostic accuracy of exhaled breath biomarker scores in cardiorespiratory disease subgroups A multinomial regression model with elastic net regularization was fitted to the matrix of 101 exhaled biomarkers using 100 iterations of 10-fold cross-validation. A linear combination of the most stable features from the multinomial regression models fitted to the 101 biomarkers formed a set of scores to predict the probability of belonging to different disease groups (acute asthma, acute COPD, pneumonia, heart failure, or healthy volunteers). Details of the median exhaled VOC scores and their distribution across disease subgroups are shown in Figure 12.
[0161] In the pooled cohort (n-277), the overall classification accuracy using all five biomarker scores was 0.722, 95% CI (0.6653-0.774) (Figure 14). The balanced accuracy for acute asthma was 0.8274, acute COPD was 0.7751, heart failure was 0.7967, community-acquired pneumonia was 0.7935, and healthy controls was 0.9274.
[0162] Discussion In this pragmatic acute treatment study, the efficacy of exhaled breath biomarker profiling in highly critically ill patients presenting with acute cardiorespiratory breathlessness was evaluated. The inventors observed that robust and validated sampling of alveolar exhaled breath using GC×GC-MS combined with GC×GC-MS biomarker characterization demonstrated high diagnostic accuracy for acute cardiorespiratory exacerbations. A putative biomarker risk score from a subset of exhaled breath VOC biomarkers that stratifies cardiorespiratory exacerbation subtypes and warrants validation in replication studies has also been identified. Furthermore, several classes of VOCS have been identified that are highly correlated and selectively increased or suppressed in acute illness (including subgroups) compared to healthy, providing potential insights into the widespread dysregulation of the metabolome in acute cardiorespiratory exacerbations.
[0163] This study is the first attempt to characterize exhaled VOCs in a large cohort with severe cardiopulmonary exacerbations, and the results position this study as a proof-of-concept for the use of breathomics in acute clinical settings.
[0164] The analytical method described here is underpinned by a robust biomarker development protocol using TD-GCxGC-FID / MS, essential for the standardization and integration of breath analysis in large-scale translational studies. Several potential confounding factors, including batch variation, were addressed in detail (SI). Furthermore, biomarker quantification of the 101 modeled VOCs followed the recommendations of the Metabolomics Standards Initiative (MSI), with 58 compounds identified against pure and traceable standards (Level I) and 21 with putative identity based on mass spectral and retention index library matches (Level 2; Figure 16). Markers that appear to localize to individual cardiorespiratory states could be readily visualized (Figure 5).
[0165] The identification of hydrocarbons and carbonyls as major chemical classes was consistent with current mechanistic understanding, which postulates them as chemical end points of lipid peroxidation, a consequence of oxidative stress during inflammation. Aldehydes such as nonanal, decanal, and hexanal predict asthma, while ketones include 2-pentanone (asthma), cyclohexanone (pneumonia), and 2,3-butanedione (COPD). Individual hydrocarbons such as 2,4- and 2,2-dimethylpentane; 2-methylbutane; 4-methyldecane; 5-methylnonane, and isoprene predict pneumonia and heart failure. Sulfur-containing VOCs such as 3-methylthiophene, arylmethyl sulfide, and carbonyl sulfide (found to predict COPD) are associated with bacterial metabolism and are thought to originate from the gut and possibly as a result of radiation injury. 2,3-butanedione also predicts COPD.
[0166] Not all compounds were considered to be endogenous VOCs, 27 of which were attributed to contamination from personal care products such as cosmetics Figure 16. Eleven of the features predicting the control group were assigned to either fragrances (e.g., alpha-isomethyl ionone) or waxy long-chain chemicals used in cosmetics as emollients and surfactants (e.g., stearyl vinyl ether and isopropyl myristate). These may have been captured in the breath samples because the sorbent tubes were close to the patients' faces.
[0167] Co-expression and enrichment analysis of the Leuven cluster on the correlation graph (Functional Enrichment Analysis section - Tables 3-6) revealed a set of highly correlated metabolites that were significantly enriched in specific disease groups. Comparison of the Leuven cluster with the metabolite set identified using the previously described method showed a strong overlap (Figure 5A and 5B). Metabolites enriched in heart failure were highly correlated C-terminally expressed metabolites with high chemical similarity. 5-7 Hydrocarbons and C 3-5 The cluster was that of carbonyls (based on the Tanimoto coefficients determined in Methods and Figure 11). The cluster included 2,4- and 2,2-dimethylpentane, 2-methylbutane, 2-methyl-1,3-butadiene (isoprene), 3-methylpentane, hexane, and cyclohexane.
[0168] This analysis also revealed that another set of highly correlated aldehydes (nonanal, decanal, undecanal, and methyldecanal isomers) were lower in asthma exacerbations compared with COPD and pneumonia exacerbations. Depletion of VOCs during in vitro experiments has been reported as a result of metabolic activity by immune cells, although the association here is tentative and should be interpreted with caution due to previously observed correlations between inhaled and exhaled air concentrations of these compounds (median Spearman rank = 0.60).
[0169] In conclusion, we have performed an acute procedure volatile exhaled breath biomarker study using robust clinical and analytical techniques and identified high diagnostic sensitivity and specificity of biomarkers in acute cardiorespiratory disease, together with robust biomarker identification and mechanistic associations that warrant further metabolomic phenotyping approaches in acute cardiopulmonary disease exacerbations.
Claims
1. A method for predicting the presence or absence of a cardiopulmonary disease in a subject, comprising: detecting the presence of one or more cardiopulmonary disease-VOC biomarkers in a breath sample from the subject; predicting that the subject has a cardiopulmonary disease if one or more VOC biomarkers are present in the sample A method as described above.
2. A method for predicting the presence or absence of a cardiopulmonary disease in a subject, the method comprising detecting the presence of one or more cardiopulmonary disease-VOC biomarkers in a breath sample from the subject, wherein the presence of one or more of the VOC biomarkers in the sample predicts that the subject has a cardiopulmonary disease, and subsequently, the subject may be administered a therapeutic agent for treating the cardiopulmonary disease A method as described above.
3. A method for selecting a subject for treatment with a therapeutic agent or composition for a cardiopulmonary disease, comprising: detecting the presence of one or more cardiopulmonary disease-VOC biomarkers in a breath sample from the subject, wherein the presence of one or more of the VOC biomarkers in the sample suggests that the subject has a cardiopulmonary disease; and selecting the subject for treatment with a therapeutic agent or composition for the cardiopulmonary disease A method as described above.
4. A method for determining whether a therapeutic agent or composition effectively treats a subject's cardiopulmonary disease, comprising: measuring the concentration of one or more cardiopulmonary disease VOC biomarkers in a test sample exhaled by the subject; and comparing the concentration of at least one or more VOCs in the test sample with the concentration in a reference sample wherein, if the concentration of one or more VOC biomarkers in the test sample is lower (or at least decreased) compared to the concentration in the reference sample, it indicates that the therapeutic agent or composition effectively treats the subject's cardiopulmonary disease. A method as described above.
5. The method according to claim 4, wherein the concentration of the VOC biomarker in the test sample is at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% lower (or at least decreased) compared to the concentration in the reference sample.
6. The method according to any one of claims 1 to 5, wherein the subject is experiencing shortness of breath.
7. The method according to any one of claims 1 to 5, which uses two-dimensional gas chromatography combined with mass spectrometry to detect the presence of one or more VOC biomarkers in a sample.
8. The method according to any one of claims 1 to 5, wherein the cardiopulmonary disease is one or more diseases selected from the group including asthma, COPD, heart failure, and pneumonia.
9. The method according to any one of claims 1 to 5, wherein one or more cardiopulmonary disease-VOC biomarkers are one or more selected from FIG.
16.
10. One or more cardiopulmonary disease-VOC biomarkers are as follows: Hexane; Octane; Tetradecane; 2,3-Butanedione; Hexanal; 2-Methyl-2-propenal; 1-Hexadecanol; 2-Methyl-1,3-dioxolane; Limonene; Eucalyptol; Menthone; p-Menta-1,4 / 8-diene; 3-Carene; Beta-farnesene; Sesquiterpenoid; Xylene; 2,3-Dimethylnaphthalene; Carbonyl sulfide; 4-Cyanocyclohexene; Methenamine; Dichloromethane; N,N-Dimethyl-1-nonanamine; Alkenyl hexanoic acid ester The method according to any one of claims 1 to 5, which is one or more selections from the above.
11. One or more cardiopulmonary disease-VOC biomarkers are one or more asthma-VOC biomarkers; one or more COPD-VOC biomarkers; one or more heart failure-VOC biomarkers; and / or one or more pneumonia-VOC biomarkers The method according to any one of claims 1 to 5.
12. One or more asthma-VOC biomarkers are as follows: 3-Methylpentane; 2-Methylnonane; Decane; 1-Nonene; Methyldecanal isomer; Undecanal; 3-Methyl-benzaldehyde; 2-Ethylhexanol; Tetrahydrofuran; 1,4-Dioxane; β-Bisabolene; and N,N-Dimethyl-1-dodecanamine The method according to claim 11, which is one or more selections from the above.
13. One or more COPD-VOC biomarkers are as follows: Nonane; 4-Methylundecane; 1-Decanol; Menthol; Camphene; Galaxolide; 3-Methylthiophene; and N,N-Dimethyl-1-dodecanamine The method according to claim 11, which is one or more selections from the above.
14. One or more heart failure-VOC biomarkers are as follows: Undecane; cyclohexene; butanal; 2-methyl-2-propenal; tridecanal; ethyl acetate; 1,3-dioxolane; betamylcene; ethylbenzene; and decyl isobutyl ether The method according to claim 11, which is one or more selections from the following:
15. One or more pneumonia-VOC biomarkers are as follows: 2,6-dimethyloctane, dimethylundecane isomer; 1-decene; 3-buten-2-one (methyl vinyl ketone); 1-(methylthio)-1-propene; 1-methylthiopropane; and one or more selections from dodecyl acrylate, the method according to claim 11.