Ai-based analysis of breath sample using chromatography
Patent Information
- Application Number
- CA3323539
- Authority / Receiving Office
- CA · CA
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-12
- Filing Date
- 2025-03-12
- Publication Date
- 2025-09-18
AI Technical Summary
Existing methods for analyzing exhaled breath biomarkers for early disease detection are hindered by the need for well-defined chromatographic peaks and suffer from cross-interferences, limiting their practical implementation and precision.
A method combining chromatography with a trained machine learning model to analyze patterns in chromatograms or graphs of breath sample components, allowing for the inference of pathologies without requiring well-separated biomarkers, using simpler detectors and reducing analysis time.
Enhances the precision and speed of breath analysis by identifying patterns in chromatograms or graphs, enabling accurate detection of pathologies without the need for precise peak separation, and facilitating cost-effective field deployment.
Abstract
Description
AI-BASED ANALYSIS OF BREATH SAMPLE USING CHROMATOGRAPHYCROSS-REFERENCE TO RELATED APPLICATION
[0000] This application claims the benefit of and priority to United States Provisional Patent Application No. 63 / 564,046, filed March 12, 2024, and entitled “AI-BASED ANALYSIS OF BREATH SAMPLE USING CHROMATOGRAPHY”, the disclosure of which is hereby incorporated by reference in its entirety.TECHNICAL FIELD
[0001] The technical field generally relates to methods and systems using Al for processing chromatography samples and more particularly concerns a method of analyzing a breath sample.BACKGROUND
[0002] The analysis of biomarkers in exhaled breath is recognized as a technology of the future to detect pathologies such as cancer in their very early stages, known as stage 0. The exhaled breath contains well over 2000 biomarkers at trace level. To properly identify a disease, multiple biomarkers must typically be analyzed, usually more than 5. The concentration of those biomarkers as well as their ratio are considered very important for most traditional analytical methods developed for exhaled breath analysis. For over 30 years, researchers have been trying to characterize all biomarkers and associate them, or a group of them, to a specific pathology, a very difficult task which has greatly reduced the pace at which such technology is implemented in the real world, for real applications.
[0003] Gas Chromatography is a very powerful tool that is used to separate the constituents of various types of samples, including breath samples. A chromatographic column separates the biomarkers based on various properties such as molecule dimension and polarity. The result is a chromatogram. A chromatogram is the representation of the chromatograph detector signal as afunction of time. The detector intensity is used to generate chromatographic peaks.A peak is associated with one molecule or biomarker.
[0004] During the past decade, artificial intelligence (Al) has been implemented in various fields. Al is a very power tool to identify patterns. In the field of exhale breath analysis, Al has been used to analyze groups of specific and well quantified biomarkers. For example, traditional chromatography, i.e. having a chromatogram with well-defined chromatographic peaks that can be quantified, has been used in combination with Al to identify patterns within well quantified molecules. Another example is the use of Al with a so-called eNose. An eNose is typically made of an array of solid-state sensors that respond to specific molecules or molecular groups. The Al in those cases is used to identify patterns and associate them with a pathology. The problem is the same as with traditional chromatography: biomarkers must be well known. The solid-state sensors used in an eNose are known to suffer from cross interferences which makes them difficult to use implement in practice. In addition, an eNose only provide signals related to specific molecules or molecular groups, while chromatography can measure 1000s of molecules in one analysis providing more information and opportunities for future reprocessing to improve exhaled breath analysis precision and specificity.
[0005] There remains a need in the art for improved methods and system for breath analysis.SUMMARY
[0006] In accordance with an aspect, a method for analysis of a breath sample is provided. The method comprises using a chromatography system comprising a chromatography column and one or more detectors, obtaining at least one chromatogram representative of an elution of components of the breath sample from the chromatography column over time, each of the at least one chromatogram defining a 2D image of a plurality of peaks as a function of time, and running a trained machine learning model at least on the at least one chromatogram to infer a presence of a pathology in the breath sample based on patterns formed by thepeaks in the at least one chromatogram. The trained machine learning model has been trained with a dataset comprising a plurality of training pairs, each of the plurality of training pairs comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
[0007] In accordance with another aspect, a method for analysis of a breath sample is provided. The method comprises using an analytical system comprising one or more detectors, obtaining at least one graph representing a distribution of components of the breath sample as a 2D image of a plurality of peaks, and running a trained machine learning model at least on the at least one graph to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one graph. The trained machine learning model has been trained with a dataset comprising a training pairs, each comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
[0008] In accordance with a further aspect, a system for analysis of a breath sample is provided. The system comprises a chromatography system comprising a chromatography column and one or more detectors, configured to obtain at least one chromatogram representative of an elution of components of the breath sample from the chromatography column over time, each of the at least one chromatogram defining a 2D image of a plurality of peaks as a function of time, and a trained machine learning model configured to run at least on the at least one chromatogram to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one chromatogram. The trained machine learning model has been trained with a dataset comprising a plurality of training pairs, each of the plurality of training pairs comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
[0009] In accordance with yet another aspect, a system for analysis of a breath sample is provided. The system comprises an analytical system comprising one or more detectors configured to obtain at least one graph representing a distribution of components of the breath sample as a 2D image of a plurality of peaks, and a trained machine learning model configured to run at least on the at least one graph to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one graph. The trained machine learning model has been trained with a dataset comprising a training pairs, each comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
[0010] In accordance with yet a further aspect, a non-transitory computer- readable medium is provided. The computer-readable medium has instructions stored thereon which, when executed by one or more processors, cause the one or more processors to, using a chromatography system comprising a chromatography column and one or more detectors, obtain at least one chromatogram representative of an elution of components of a breath sample from the chromatography column over time, each of the at least one chromatogram defining a 2D image of a plurality of peaks as a function of time, and run a trained machine learning model at least on the at least one chromatogram to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one chromatogram. The trained machine learning model has been trained with a dataset comprising a plurality of training pairs, each of the plurality of training pairs comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
[0011] In accordance with yet another aspect, a non-transitory computer-readable medium is provided. The computer-readable medium has instructions stored thereon which, when executed by one or more processors, cause the one or more processors to, using an analytical system comprising one or more detectors, obtain at least one graph representing a distribution of components of a breath sampleas a 2D image of a plurality of peaks, and run a trained machine learning model at least on the at least one graph to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one graph. The trained machine learning model has been trained with a dataset comprising a training pairs, each comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] For a better understanding of the embodiments described herein and to show more clearly how they may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings which show at least one exemplary embodiment.
[0013] FIG. 1 is a schematic representation of a gas chromatography system, in accordance with an embodiment.
[0014] FIG. 2 is a flowchart of a method for analysis of a breath sample, in accordance with an embodiment.
[0015] FIG. 3A (PRIOR ART) is an illustration of a chromatogram known in the prior art, for comparison purpose; FIG. 3B is an illustration of a chromatogram for analysis of a breath sample, in accordance with an embodiment.
[0016] FIG. 4 is an illustration of a neural network for analysis of a breath sample, in accordance with an embodiment.
[0017] FIG. 5 is a schematic representation of a method for training a machine learning model for analysis of a breath sample, in accordance with an embodiment.
[0018] FIG. 6 is a schematic representation of a method for analysis of a breath sample, in accordance with an embodiment.
[0019] FIGs. 7A and 7B are illustrations of chromatograms for analysis of a breath sample, in accordance with an embodiment.
[0020] FIG. 8 is a schematic representation of a method for analysis of a breath sample, in accordance with an embodiment.
[0021] FIG. 9 shows illustrations of chromatograms for analysis of a breath sample, in accordance with an embodiment.
[0022] FIG. 10 is a flowchart of a method for analysis of a breath sample known in the prior art, for comparison purpose.
[0023] FIGs. 11 and 12 (PRIOR ART) are illustrations of methods for analysis of a breath sample known in the prior art, for comparison purpose.
[0024] FIGs. 13 to 17 are flowcharts of methods for analysis of a breath sample, in accordance with an embodiment and various use cases.
[0025] FIG. 18 is a schematic representation of a method for training a neural network, in accordance with an embodiment and a use case.
[0026] FIG. 19 is a flowchart of a method for using a neural network, in accordance with an embodiment and a use case.
[0027] FIG. 20A (PRIOR ART) is a flowchart of a method to detect cancer biomarkers in a breath sample known in the prior art, for comparison purpose; FIG. 20B is a flowchart of a method to detect cancer biomarkers in a breath sample, in accordance with an embodiment.
[0028] FIGs. 21 A and 21 B are illustrations of chromatograms for analysis of a breath sample, in accordance with an embodiment and a use case.DETAILED DESCRIPTION
[0029] In the present description, similar features in the drawings have been given similar reference numerals. To avoid cluttering certain figures, some elements maynot be indicated if they were already identified in a preceding figure. It should also be understood that the elements of the drawings are not necessarily depicted to scale, since emphasis is placed on clearly illustrating the elements and structures of the present embodiments. Positional descriptors indicating the location and / or orientation of one element with respect to another element are used herein for ease and clarity of description. Unless otherwise indicated, these positional descriptors should be taken in the context of the figures and should not be considered limiting. It is appreciated that such spatially relative terms are intended to encompass different orientations in the use or operation of the present embodiments, in addition to the orientations exemplified in the figures. Furthermore, when a first element is referred to as being “on”, “above”, “below”, “over”, or “under” a second element, the first element can be either directly or indirectly on, above, below, over, or under the second element, respectively, such that one or multiple intervening elements may be disposed between the first element and the second element.
[0030] The terms “a”, “an”, and “one” are defined herein to mean “at least one”, that is, these terms do not exclude a plural number of elements, unless stated otherwise.
[0031] The term “or” is defined herein to mean “and / or”, unless stated otherwise.
[0032] The expressions “at least one of A, B, and C” and “one or more of A, B, and C”, and variants thereof, are understood to include A alone, B alone, and C alone, as well as any combination of A, B, and C.
[0033] Terms such as “substantially”, “generally”, and “about”, which modify a value, condition, or characteristic of a feature of an exemplary embodiment, should be understood to mean that the value, condition, or characteristic is defined within tolerances that are acceptable for the proper operation of this exemplary embodiment for its intended application or that fall within an acceptable range of experimental error. In particular, the term “about” generally refers to a range of numbers that one skilled in the art would consider equivalent to the stated value(e.g., having the same or an equivalent function or result). In some instances, the term “about” means a variation of ±10% of the stated value. It is noted that all numeric values used herein are assumed to be modified by the term “about”, unless stated otherwise. The term “between” as used herein to refer to a range of numbers or values defined by endpoints is intended to include both endpoints, unless stated otherwise.
[0034] The term “based on” as used herein is intended to mean “based at least in part on”, whether directly or indirectly, and to encompass both “based solely on” and “based partly on”. In particular, the term “based on” may also be understood as meaning “depending on”, “representative of”, “indicative of”, “associated with”, and the like.
[0035] The terms “match”, “matching”, and “matched” refer herein to a condition in which two elements are either the same or within some specified tolerance of each other. That is, these terms are meant to encompass not only “exactly” or “identically” matching the two elements but also “substantially”, “approximately”, or “subjectively” matching the two elements, as well as providing a higher or best match among a plurality of matching possibilities.
[0036] The terms “connected” and “coupled”, and derivatives and variants thereof, refer herein to any connection or coupling, either direct or indirect, between two or more elements, unless stated otherwise. For example, the connection or coupling between elements may be mechanical, optical, electrical, magnetic, thermal, chemical, logical, fluidic, operational, or any combination thereof.
[0037] The term “concurrently” refers herein to two or more processes that occur during coincident or overlapping time periods. The term “concurrently” does not necessarily imply complete synchronicity and encompasses various scenarios including time-coincident or simultaneous occurrence of two processes; occurrence of a first process that both begins and ends during the duration of a second process; and occurrence of a first process that begins during the duration of a second process, but ends after the completion of the second process.
[0038] The present description generally relates to methods for analysis of a breath sample containing different biomarkers representative of a condition of a patient.
[0039] According to one aspect, the present description concerns a method for analysis of a breath sample using a chromatography system with a chromatography column and one or more detectors, the method involves obtaining at least one chromatogram representative of an elution of components of the breath sample from the chromatography column over time. Each chromatogram defines a 2D image of a plurality of peaks as a function of time. The method further includes running a trained machine learning model on the chromatogram(s) to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one chromatogram.
[0040] The trained machine learning model has preferably been trained with a dataset comprising a plurality of training pairs, each of the plurality of training pairs comprising at least one training chromatogram, each of the at least one training chromatogram defining a 2D training image, and at least one label, each of the at least one label corresponding to a training pathology.
[0041] As will be clear from a reading of the description below, the method presented herein makes use of Al to approach exhaled breath analysis in a completely new way and provides a novel approach to interpreting chromatograms. Advantageously, this approach does not require biomarkers to be well separated in the signal eluting from the chromatography column. They only need to generate a signal that is noticeable on the chromatogram, as opposed to standard chromatography relies on properly defined and quantifiable peaks.
[0042] Although the present description generally applies the present method to information obtained from a chromatography system, it will be readily understood that the method may also be used with other types of analytical systems that generate data with sufficient resolution to show the presence of multiple biomarkers. By way of example, the analytical system may be a FTIR (Fourier-Transform Infrared Spectroscopy) system, which provides as output a spectrum representative of the absorption of emission of infrared light by the breath sample. The spectrum is represented as a graph which defines a 2D image of a plurality of peaks. This 2D image may be processed in the same manner as described below for the chromatograms obtained in the context of chromatography.
[0043] The term “chromatography” refers herein to an analytical or process technique for separating a sample or mixture into its individual components and analyzing qualitatively and quantitatively the separated sample components. In most chromatography applications, the sample is transported in a carrier fluid to form a mobile phase. Depending on whether the mobile phase is a gas or a liquid, chromatography can be classified into two main branches: gas chromatography (GC) and liquid chromatography (LC), both of which can be used to implement the present techniques. The mobile phase is then carried through a stationary phase, which is located in a column or another separation device. The mobile and stationary phases may be selected so that the components of the sample transported in the mobile phase exhibit different interaction strengths with the stationary phase. This leads to different sample components having different retention times through the system, where the sample components that are strongly interacting with the stationary phase move more slowly with the flow of the mobile phase and elute from the column later than the sample components that are weakly interacting with the stationary phase. As the sample components separate, they elute from the column and enter a detector. The detector is configured to generate an electrical signal whenever the presence of a sample component is detected. Typically, the magnitude of the signal is proportional to the concentration level of the detected component. The measurement data can be processed by a computer to obtain a chromatogram, which is a time series of peaks representing the sample components as they elute from the column. The retention time of each peak is indicative of the composition of the corresponding eluting component, while the peak height or area conveys information on the amount or concentration of the eluting component.
[0044] The term “sample” refers herein to any substance known, expected, or suspected of containing analytes. Samples can be broadly classified as organic, inorganic, or biological, and can be further subdivided into solids, semi-solids (e.g., gels, creams, pastes, suspensions, colloids), liquids, and gases. Chromatographic samples generally receive some type of pre-treatment or conditioning prior to chromatography analysis. Samples can include a mixture of analytes and nonanalytes. The term “analyte” is intended to refer to any sample component of interest that can be analyzed by chromatography, while the term “non-analyte” is intended to refer to any sample component for which chromatography analysis is not of interest in a given application. Non-limiting examples of non-analytes can include, to name a few, water, oils, solvents, and other media in which analytes may be found, as well as impurities and contaminants. It is noted that in some instances, the term “sample components” may be used interchangeably with the term “analytes” to refer to components of interest of a sample.
[0045] The present techniques have potential use in GC-based breath gas analysis for the detection and diagnosis of diseases and disorders. Breath gas analysis is a non-invasive tool for medical research and diagnosis, which can be used to gain clinical information on the physiological state of an individual. Exhaled breath is mainly composed of nitrogen, oxygen, carbon dioxide, water vapour, and argon, which are generally non-analytes, along with trace amounts of volatile organic compounds (VOCs), among which some may be diagnostically useful analytes. Typically, breath gas analysis involves the identification and quantification of analytes that provide biomarkers indicative of various pathologies and conditions. Non-limiting examples of such pathologies and conditions can include, to name a few, cancer, respiratory, pulmonary, kidney, and liver diseases, diabetes, alcohol intoxication, organ rejection, sleep apnea, and mental and physical stress.
[0046] Various implementations of the present techniques will now be described with reference to the figures. It is noted that the figures generally depict fluid flows with solid lines and communication links with dashed lines.
[0047] FIG. 1 is a schematic representation of a possible embodiment of a gas chromatography (GC) system 100 that can be used to implement the present techniques. The GC system 100 allows for the separation (or partial separation) of a vaporized or gaseous sample into its components, by passing a mobile phase carrying the gas sample through a stationary phase, and the subsequent detection and analysis of the separated components. In the illustrated embodiment, the gas sample is an exhaled breath sample, although various other types of samples, including gases, vaporized liquids, and vaporized solids, can be analyzed in other embodiments. It is also appreciated that although the illustrated embodiment is directed to a GC system, the present techniques may be implemented in any suitable chromatography system, including LC systems, such as high-performance liquid chromatography (HPLC) systems.
[0048] The GC system 100 of FIG. 1 generally includes a sample handling unit 102, a carrier gas unit 104, a chromatographic separation unit 106, a detection unit 108, and a control and processing unit 110. More detail regarding the structure and operation of these units and other possible components of the GC system 100 will be provided below. It is appreciated that FIG. 1 is a simplified schematic representation that illustrates a number of basic components of the GC system 100, such that additional features and components that may be useful or necessary for the practical operation of the GC system 100 may not be specifically depicted. Non-limiting examples of such additional features and components can include, to name a few, purge lines to remove dead volumes, pressure and flow regulators, restrictors, and other standard hardware and equipment. It is also appreciated that the theory, instrumentation, operation, and application of chromatography systems and methods are generally known in the art and need not be described in detail herein other than to facilitate an understanding of the present techniques.
[0049] The sample handling unit 102 includes the instrumentation of the GC system 100 that is configured for handling the sample to analyze, which can include steps such as collecting or receiving the sample; processing the sample to make it suitable for GC analysis; and dosing and injecting the processed sampleas part of a mobile phase into the chromatographic separation unit 106. As can be appreciated, the sample handling unit 102 can include various hardware components, such as inlets, outlets, flow lines, containers, chambers, sample loops and traps, flow directing and regulating equipment (e.g., pumps and valves, flow metering equipment, filters), and the like, all of which are generally known in the art and need not be described in detail herein.
[0050] In FIG. 1 , the sample handling unit 102 generally includes a sample collector 112, a sample conditioner 114, and a sample injector 116. It is appreciated that although the sample handling unit 102 in FIG. 1 is described as including modules performing specific functions (i.e., collection, conditioning, and injection), other embodiments may include modules that combine and add to the functions of the modules of the sample handling unit 102 depicted in FIG. 1.
[0051] In some implementations, the sample collector 112 may include any device or combination of devices configured for collecting or sampling exhaled breath from a subject 118, which can for example involve exhalation through a face mask or into a bag.
[0052] The sample conditioner 114 may include any device or combination of devices configured for preparing, treating, or otherwise conditioning the collected sample into a gas sample suitable for injection as part of a mobile phase into the chromatographic separation unit 106. It is appreciated that the sample conditioner 114 can be configured to perform a variety of processing steps. Non-limiting examples of such steps can include, to name a few, steps of controlling the temperature, pressure, concentration, and / or flow of the sample, as well as various other pretreatment steps, such as sample accumulation, storage, filtering, division, purification, extraction, drying, vaporization, derivatization, enhancement, mixing, and any combination thereof. Finally, in some implementations the GC system may omit any form of sample conditioning and directly input the collected sample into the sample injector 116.
[0053] The sample injector 116 can include any device or combination of devices configured for injecting the gas sample into the chromatographic separation unit 106. In FIG. 1 , the gas sample is introduced into the carrier gas flow supplied by the carrier gas unit 104, and the resulting sample-carrier gas mixture is flown, as the mobile phase, into the chromatographic separation unit 106. The sample injector 116 may include a syringe, which can be manually operated or part of an autosampler, or other fluid delivery devices configured for injecting the gas sample into the stream of carrier gas supplied by the carrier gas unit 104. In FIG. 1 , the sample injector 116 includes a multiport switching valve 120, although other fluid directing devices for combining and directing fluid flows may be provided in other embodiments. In FIG. 1 , the multiport switching valve 120 is a six-port switching valve having a first inlet port 122a connected to the carrier gas unit 104, a second inlet port 122b connected to the sample conditioner 114 (e.g., via a syringe), a first outlet port 122c connected to the chromatographic separation unit 106, a second outlet port 122d defining a sample vent, and two ports 122e, 122f defining a sample loop 124. In other embodiments, the sample loop 124 may be replaced by a sample trap configured for sample concentration. It is noted that in FIG. 1 , the multiport switching valve 120 is depicted in a sample injection configuration, in which the connections between the inlet and outlet ports are depicted by solid lines. It is appreciated that the multiport switching valve 120 can be switched to a sample loading configuration, in which the connections between the inlet and outlet ports are depicted by dotted lines. The structure and operation of multiport switching valves used for sample injection in GC applications are known in the art and need not be described in detail herein. It is also appreciated that the present techniques can employ any suitable sample injection method, which can be performed on an automated, semi-automated, or manual basis.
[0054] The carrier gas unit 104 can include any device or combination of devices configured for supplying a flow of carrier gas for providing a suitable mobile phase for conveying the gas sample into and through the chromatographic separation unit 106. The carrier gas unit 104 can include a carrier gas source (e.g., a gas storage tank or cylinder), carrier gas supply lines (e.g., conduits, such as tubes orpipes) to convey the carrier gas between the carrier gas source and the chromatographic separation unit 106 via the sample handling unit 102, and flow regulators (e.g., pumps, valves, and restrictors) to control the carrier gas flow rate and pressure. The carrier gas is introduced with the gas sample upstream of the chromatographic separation unit 106, typically at or near the sample injector 116, for example, via a multiport switching valve 120, such as depicted in FIG. 1 (e.g., via the first inlet port 122a). The carrier gas may be any gas capable of providing a suitable mobile phase for carrying the gas sample, non-limiting examples of which can include, to name a few, helium, nitrogen, argon, air, oxygen, and hydrogen.
[0055] The chromatographic separation unit 106 may include a GC column 126 or another chromatographic device or combination of devices able to separate the gas sample into its constituents. Non-limiting examples of GC columns include packed columns and capillary columns. The GC column 126 has a column inlet 128 and a column outlet 130. In the illustrated implementation, the column inlet 128 is fluidly connected to the sample handling unit 102 and the carrier gas unit 104 via the sample injector 116 (e.g., via the first outlet port 122c of the multiport switching valve 120). The column outlet 130 is fluidly connected to the detection unit 108. The GC column 126 may be housed in or otherwise associated to a thermally controlled GC oven 132. The GC oven 132 can have any appropriate configuration for maintaining the GC column 126 at a selected temperature or for varying the temperature of the GC column 126 according to a selected temperature profile. The chromatographic separation unit 106 contains a stationary phase suitable for GC, which can include particles or films deposited on the inner surface of the GC column 126. The composition of the stationary phase may be selected so that different sample components in the gas sample exhibit different interaction strengths with the stationary phase, leading to different retention times. Thus, as the mobile phase flows through the GC column 126, the gas sample gradually separates into discrete sample components. The separated sample components eluting at the column outlet 130 are received and detected by the detection unit 108.
[0056] The detection unit 108 may include any appropriate detector or combination of detectors configured for providing qualitative and / or quantitative measurement of the separated sample components eluting from the chromatographic separation unit 106. Generally, the detection unit 108 is configured to respond to a property of the separated sample components, generate electrical detection signals based on the measured property, convert the generated electrical detection signals into digital detection signals, and supply the digital detection signals to the control and processing unit 110 for analysis, display, and / or storage. Various types of chromatography detectors exist and can be classified in different ways, including destructive or non-destructive, selective or universal, and concentration-sensitive or mass-sensitive detectors. Non-limiting examples of chromatography detectors include a flame ionization detector (FID), a thermal conductivity detector (TCD), an electron capture detector (ECD), a flame thermionic detector (FTD), a flame photometric detectors (FPD), an atomic emission detector (AED), an optical emission spectroscopy (OES) detector, a nitrogen phosphorus detector (NPD), a mass spectrometer (MS), electrolytic conductivity detector (ELCD), a plasma emission detector (PED), an enhanced plasma discharge (Epd) detector, a UV detector, a fluorescence detector, and a photoionization detector (PID). In some embodiments, the detection unit 108 may include a combination of two or more different types of detectors. For example, in some embodiments, the detection unit 108 may be integrated into an analytical instrument, such as a mass spectrometer (MS) or an ion mobility spectrometer (IMS). The choice of a suitable detector can be made based on a number of factors, such as the type and number of analytes to be detected, the type of analytical information to be derived from the detected analytes, the desired or required accuracy and precision of the analysis, cost and space considerations, and the like.
[0057] The control and processing unit 110 is configured for controlling, monitoring, and / or coordinating the functions and operations of various components of the GC system 100, such as, for example, the sample handling unit 102, the carrier gas unit 104, the chromatographic separation unit 106, and thedetection unit 108, as well as various temperature, pressure, and flow rate conditions. The control and processing unit 110 is also configured to process and / or analyze the detection signals received from the detection unit 108 to derive information about the presence and / or concentration of analytes in the gas sample under analysis. In some implementations, the control and processing unit 110 can process the detection signals into a chromatogram, which is a graphical representation of the detector response plotted as a function of retention time. The chromatogram generally provides a spectrum of peaks representing the analytes present in the sample eluting from the chromatographic separation unit 106 and into the detection unit 108 at different retention times. As further detailed below, each chromatogram defines a corresponding 2D image of a plurality of peaks as a function of time. This 2D image can be called a “breathprint”.
[0058] In some implementations, the control and processing unit 110 is further configured to implement artificial intelligence algorithms to identify patterns in one or more chromatogram and associate the patterns with one or more pathology. In some embodiments, a trained machine learning model is provided to that end. The machine learning model can be trained to accept one or more 2D image(s) each corresponding to a chromatogram, i.e., one or more breathprint(s). In some embodiments, the machine learning model is trained to accept as further input data characterizing the individual, e.g., data stored in and provided by a records system such as an electronic health records system, for instance including personal information such as the age, the weight and / or the height of the individual, and / or medical information related to the individual such as current and / or historical vital signs, symptoms, body fluid, e.g., blood, analysis results and / or known medical conditions, e.g., whether the individual has received a diabetes mellitus diagnostic.
[0059] Using a trained model such as a deep learning model has the benefit that such model can be optimized and updated as new data becomes available. In some embodiments, one model can be trained for each condition, e.g., pathology, of interest. The end result can be a bounded or unbounded, e.g., real value, indicative of whether the condition of interest is present in the individual whoprovided the breath sample, or a probability score within a set range such as [0,1 ], or a pair of probability scores within a set range such as [0,1 ], optionally defined such that the sum of the two scores is 1 , with a first score indicating a probability that the condition of interest is present, i.e. , is identifiable in the data provided as input to the model, and the second score indicating a probability that the condition of interest is not present. In some embodiments, one, some or all models can be trained to target more than one condition of interest, e.g., a subset of the conditions of interest or all conditions of interest. The end result can be a vector of bounded or unbounded, e.g., real value, indicative of whether each condition of interest associated with the model are present in the individual who provided the breath sample, or a vector of probability scores within a set range such as [0, 1 ], optionally defined such that the sum of all the scores is 1 , defining a set of probabilities that the input is indicative of each condition of interest associated with the model.
[0060] The trained machine learning model can correspond to any suitable type of classifier that contains layers to process image inputs, such as a neural network model. With reference to FIG. 4, in some embodiments, the model corresponds to a model of a convolutional neural network (CNN) 400. The 2D image(s) 320 can be converted to a tensor, e.g., a bidimensional or a tridimensional tensor, as described below with respect to method step 207. As an example only, CNN 400 can comprise any suitable number of convolutional layers 411 , 413, 415, 417, such that the input tensor corresponding to the 2D image(s) 320 serves as the input of the first convolutional layer 411 , each layer being configured to apply filters or kernels of a suitable size, for instance 3 x 3, with appropriate padding and stride, for instances 0 and 1 , then to apply a suitable activation function to each output value, for instance the rectified linear unit (ReLu) function. It can be appreciated that different numbers of convolutional layers and filters, different filter sizes, different padding sizes, different stride sizes and different activation functions can be used. In some embodiments, a batch normalization layer can be applied between any two convolutional layers. In some embodiments, pooling layers 421 , 423, 425 can be applied between any two convolutional layers, implementing a suitable pooling function, for instance max pooling. The output of the lastconvolutional layer 417 can serve as the input of the first of any suitable number of fully connected linear layers 431 , 433, configured such that the last linear layer 433 outputs a number of values corresponding to the number of classes considered by the classifier. As an example, the last linear layer 433 can output two values, corresponding to probabilities that the input 320, respectively, is indicative of a condition of interest and is not indicative of the condition of interest. The output of the last linear layer 433 can serve as the input of a softmax layer 441 , which applies a softmax function to the input values such that it outputs the same number of values and that the sum of the output values is 1 .
[0061] In the present exemplary architecture, the convolutional layers 411-417 can be said to define an encoder which embeds the input tensor 320 into a vector with a certain dimensionality, and the fully connected layers 431-433 can be said to define a classifier which takes the vector embedding as an input and provides a classification, e.g., a probability or a set of probabilities, as an output. In some embodiments, an early data fusion approach is used. As an example, a tridimensional tensor corresponding to a plurality of breathprints from the same breath sample can be fed in a CNN such as CNN 400, with each breathprint being treated as a distinct input channel. In some embodiments, an intermediate data fusion approach is used. As an example, a plurality of encoders can be provided such that each matrix of a plurality of matrices corresponding to a plurality of breathprints for the same breath sample can be fed as input to a distinct first convolutional layer 411 , resulting in a distinct vector embedding being output by the corresponding distinct last convolutional layers 417. In some embodiments, all the encoders are identical. In some embodiments, matrices, i.e., bidimensional tensors, can be fed to a suitable number of distinctly trained encoder. As an example, a distinct encoder can be trained to process breathprints coming from distinct types of detectors. In intermediate data fusion embodiments, the vector embeddings output by the plurality of distinct last convolutional layers 417, each having the same or a different dimensionality, can be aggregated, for instance by concatenation or by interpolation. In some embodiments, the output of one or more last convolutional layer(s) 417 can be aggregated, e.g., concatenated orinterpolated, with a vectorial embedding corresponding to the individualcharacterizing data before being fed into a suitable first fully connected layer 431. In some embodiments, the vectorial embedding is an arbitrary embedding defined by the system implementation. As an example, each of d data can be converted to a suitable numeric format to be combined in a vector, e.g., a d-dimension vector. Alternatively or additionally, the data can be fed as input to a trained embedding model which outputs a vector of a predetermined dimensionality. The output of the embedding model can then be aggregated with the output of one or more last convolutional layer(s) 417.
[0062] It can be appreciated that different neural network architectures are suitable to detecting whether tensors representing a breathprint can be used to estimate the probability that the breathprint is indicative of a condition of interest. As an example, different CNN architectures such as Visual Geometry Group networks, residual neural networks, DenseNets and MobileNets can be used to classify breathprints. As other examples, vision transformers, recurrent neural networks, capsule neural networks, generative adversarial networks and graph neural networks, can be used to classify breathprints. In some embodiments, an ensemble model including one or more of the types of models described above and / or alternative types of models is used to classify breathprints. It can further be appreciated that the neural network architectures described above are equally suitable for processing images representing of breath-related plotted data, such as spectrums.
[0063] Because biomarkers will elute at the same time, using image-based trained machine learning models simplifies exhaled breath analysis by removing the requirement to be aware of the nature and significance of the biomarkers. Additionally, because the shape of the chromatogram or portions of the chromatogram are being analyzed by the model and the peak area of the chromatogram is not being quantified, biomarkers coelution, i.e. , biomarker peaks not being properly separated from one another, does not affect the accuracy of the result. Moreover, this approach allows using much simpler chromatographicdetectors. Currently, the mass spectrometer is the detector of choice because it allows the identification of molecules based on their charge-to-mass ratio. The mass spectrometer is a powerful tool for research, but is not suitable for field deployment for general use. Interpreting the chromatogram as an image makes it possible to use simpler and less selective detectors, because it is not necessary to positively identify specific molecules. Rather, the purpose is to generate a normalized image of the exhaled breath for machine classification. This approach simplifies the implementation and greatly reduces the cost. Moreover, the imagebased approach described in the present disclosure significantly reduces the time required to perform chromatographic exhaled breath analysis. Typical approaches require properly resolving the biomarkers, which can take upward of 60 minutes when thousands of compounds are being analyzed. A faster analysis would lead to biomarker coelution, making the analysis task very complex for traditional GC software and therefore increase processing time. Complex de-convolution can be used to improve quantification, but this is complex, unreliable and prone to errors. Using the approach described herein, a faster chromatogram can be performed, even with overlapping biomarkers, because peaks need not be properly resolved to generate a noticeable signal in the image.
[0064] The control and processing unit 110 can be implemented in hardware, software, firmware, or any combination thereof, and be connected to various components of the GC system 100 via wired and / or wireless communication links to send and / or receive various types of signals, such as timing and control signals, measurement signals, and data signals. The control and processing unit 110 may be controlled by direct user input and / or by programmed instructions, and may include an operating system for controlling and managing various functions of the GC system 100. Depending on the application, the control and processing unit 110 may be fully or partly integrated with, or physically separate from, the other hardware components of the GC system 100. In FIG. 1 , the control and processing unit 110 generally includes a processor 134 and a memory 136.
[0065] The processor 134 may be able to execute computer programs, also generally known as commands, instructions, functions, processes, software codes, executables, applications, and the like. It should be noted that although the processor 134 in FIG. 1 is depicted as a single entity for illustrative purposes, the term “processor” should not be construed as being limited to a single processor, and accordingly, any known processor architecture may be used. In some implementations, the processor 134 may include a plurality of processing units. Such processing units may be physically located within the same device, or the processor 134 may represent processing functionality of a plurality of devices operating in coordination. For example, the control and processing unit 110 may include a main processor configured to provide overall control and one or more secondary processors configured for dedicated control operations or signal processing functions. Depending on the application, the processor 134 may include or be part of a computer; a microprocessor; a microcontroller; a coprocessor; a central processing unit (CPU); an image signal processor (ISP); a digital signal processor (DSP) running on a system on a chip (SoC); a single-board computer (SBC); a dedicated graphics processing unit (GPU); a special-purpose programmable logic device embodied in hardware device, for example, a field- programmable gate array (FPGA) or an application-specific integrated circuit (ASIC); a digital processor; an analog processor; a digital circuit designed to process information; an analog circuit designed to process information; a state machine; and / or other mechanisms configured to electronically process information and to operate collectively as a processor.
[0066] It is understood that the neural network(s) can be implemented using computer hardware elements, computer software elements or a combination thereof. Accordingly, the neural networks described herein can be referred to as being computer-implemented. Various computationally intensive tasks of the neural network can be carried out on one or more processors 134 (central processing units and / or graphical processing units) of one or more programmable computers. For example, and without limitation, the programmable computer may be a programmable logic unit, a mainframe computer, server, personal computer,cloud-based program or system, laptop, personal data assistant, cellular telephone, smartphone, wearable device, tablet device, virtual reality device, smart display devices such as a smart TV, set-top box, video game console, or portable video game device, among others.
[0067] The memory 136, which can also be referred to as a computer-readable storage medium, is capable of storing computer programs and other data to be retrieved by the processor 134. The terms “computer-readable storage medium” and “computer readable memory” are intended to refer herein to a non-transitory and tangible computer product that can store and communicate executable instructions for the implementation of various steps of the methods disclosed herein. The computer readable memory can be any computer data storage device or assembly of such devices, including random-access memories (RAMs); dynamic RAMs; read-only memories (ROMs); magnetic storage devices, such as hard disk drives, solid-state drives, floppy disks, and magnetic tapes; optical storage devices, such as compact discs (e.g., CDs and CDROMs), digital video discs (DVDs), and Blu-RayTM discs; flash drive memories; and / or other non- transitory memory technologies. A plurality of such storage devices may be provided, as can be appreciated by those skilled in the art. The computer readable memory may be associated with, coupled to, or included in a computer or processor configured to execute instructions contained in a computer program stored in the computer-readable memory and relating to various functions associated with the computer or processor.
[0068] The GC system 100 may also include one or more user interface devices operatively connected to the control and processing unit 110 to allow the input of commands and queries to the GC system 100, as well as present the outcomes of the commands and queries. The user interface devices can include input devices 138 (e.g., a touch screen, a keypad, a keyboard, a mouse, a switch, and the like) and output devices 140 (e.g., a display screen, a printer, visual and audible indicators and alerts, and the like).
[0069] Breath Analysis Method
[0070] Referring to FIG. 2, in according to some implementations, there is provided a method 200 for analysis of a breath sample.
[0071] The method includes using a chromatography system to obtain 206 at least one chromatogram, the chromatography system includes a chromatography column and one or more detectors. In one example, the chromatography system corresponds to the gas chromatography system 100 shown in FIG. 1 and described above, or variants thereof. It will, however, be readily understood that chromatography systems having different configurations, for example based on liquid chromatography, may be used without departing from the scope of protection.
[0072] In some implementations, the method may include preliminary steps of acquiring 202 the breath sample and injecting 204 the same into the chromatography system. For example, a breath sampling apparatus configured to collect VOCs from the breath exhaled by the subject with minimum crosscontamination from other sources may be used.
[0073] Each chromatogram is representative of the elution of components of the breath sample over time. As explained above, different biomarkers in the breath sample have different retention times in a GC column so that they are eluted separately from the column and reach the detector at different moments. The detector typically outputs an electrical signal of varying intensity over time, which is plotted as a graph. The resulting graphical representation embodies the chromatogram. Each chromatogram therefore defines a 2D image of a plurality of peaks as a function of time.
[0074] In accordance with one aspect, the present method processes the chromatogram as a 2D image rather than a spectrum to be analyzed.
[0075] In prior art chromatography methods, software is used to find peaks in the data from the detector which represent different molecules by their retention time.The retention time of each peak is typically used to identify the sample component, while the peak height or area is used to the determine the quantity of the sample component in the gas sample. When a peak is discovered, its area is calculated and processed through a transfer function known as the calibration curve. Other peak parameters, such as peak shape and peak width, can also be considered in the analysis. For this type of analysis to work well, the peaks must be well resolved from all other analytes and ideally, baseline resolved. A traditional chromatogram identifying different peaks analyzed using such a process is shown in FIG 3A (PRIOR ART).
[0076] Like certain radiological images, a chromatogram is a 2D image. This is a representation, a fingerprint, of a subject’s breath profile. By interpreting the chromatogram as a 2D image, the knowledge of the precise nature of the biomarkers present in the breath sample becomes irrelevant, as it is the shape of the peaks appearing on this image which is a representation of the metabolic health of the subject. Patterns in such images can be associated with different pathologies or other conditions as long as the corresponding biomarkers are consistently eluting at the same time for each breath sample being analyzed. FIG 3B shows the 2D image 320 corresponding to the chromatogram of FIG. 3A. In this case no labelling of peaks of axes is necessary, only the shape of the represented peaks is of interest.
[0077] In some embodiments, the method 200 may involve processing 207 the chromatogram using image recognition and / or classification techniques. By way of example, the breathprint(s) associated with the breath sample can be used to generate a data structure of a form suitable for use as input by the trained machine learning model. As another example, a matrix, also called a bidimensional tensor, can be generated, for instance by a rasterization module, from a rasterized version of each 2D image by associating a value with each matrix cell, and each matrix can be fed as input to a deep learning model. Various schemes can be used to attribute a value to each matrix cell. As an example, cells corresponding to pixels of the raster image where the plotline is visible can be associated with a value of1 and other cells can be associated with a value of 0. As another example, cells corresponding to pixels in the area under the curve in the raster image can be associated with a value of 1 and other cells can be associated with a value of 0. In both of the above examples, the matrix can be described as a bidimensional mask. As a further example, the raster image can display a smoothed curve, e.g., by using greyscale, and each cell can receive a value of 0 to indicate a black pixel, a value of 1 to indicate a white pixel, or a value greater than 0 but less than 1 to indicate a grey pixel.
[0078] Still referring to FIG. 2 and with additional reference to FIG. 5, the method further includes running 208 a trained machine learning model 570 on the chromatogram to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the chromatogram. Preferably, the trained machine learning model has been trained with a dataset 560 comprising a plurality of training pairs 550. Each training pair 550 includes one or more breathprints 541 , 543, i.e., 2D image defined by corresponding training chromatograms, which correspond to the readings obtained by a corresponding number of detectors 521 , 523 operating on GC columns 521 , 523 for a breath sample 511 obtained from an individual 510 at a given time, the individual either having one or more known condition of interest, e.g., pathology, or being known to be free of any condition of interest. Each training pair 550 also includes a label, also called a target, which is an indication of whether the individual has one or more condition of interest, and which. In some embodiments, each training pair 550 further includes data characterizing the individual, for instance, including personal information and / or medical information related to the individual as described above. In some embodiments, when additional information and / or breath sample(s) are obtained from an individual already represented in the dataset are obtained at a different point in time, the label and / or the individual-characterizing data associated with previous learning pairs associated with the same individual at previous points in time can be updated. As an example, if a pathology is diagnosed at a time t in an individual, a training pair corresponding to the same individual at a previous time t-1 can be labelled accordingly. Because chromatograms can contain signals fromthousands of VOCs, this approach is especially advantageous as it makes it possible to learn and reprocess the signals overtime. With a eNose-like approach, which takes a small number of biomarkers into account, too much information is lost and it is not possible to follow the evolution of an individual’s health or to improve the system’s accuracy as the Al learn new patterns.
[0079] Preferably, the dataset 560 will include thousands of breathprints associated conditions — or absence thereof. In some embodiments, the dataset 560 is divided into a training dataset and a validation and / or a testing dataset. The validation dataset includes training pairs that are kept out of the training dataset and of the testing dataset to evaluate the performance of the model while it is being trained. The test dataset includes training pairs that are kept out of the training dataset and of the validation dataset to evaluate the performance of the model after it has been trained. In some embodiments, as the dataset grows and is improved, so is the model improved, for instance by retraining and / or by incremental learning.
[0080] In some embodiments, the trained machine learning model 570 is a neural network model. In these embodiments, a model with the desired architecture, for instance the architecture illustrated in Figure 4, can be initialized with random parameters, e.g., weights. Tensors associated with the chromatograms in the training dataset are input in the model and a prediction is obtained. The prediction can for instance include two real numbers comprised between 0 and 1 , such that the first number is an estimation p of the probability that the input tensor corresponds to a breath sample of an individual with a condition of interest, the second number is an estimation 1 -p of the probability that the input tensor corresponds to a breath sample of an individual without the condition of interest, and the sum of the two numbers is 1 . It can be appreciated that the nature of the predicted probabilities corresponds to the nature of the possible labels used. As an example, if the labels correspond to one of n classes corresponding to n conditions, the model can be trained to output n probabilities. For each training pair of the training dataset, the predictions and the labels can be passed as argumentsto any suitable loss function, for instance a mean square error function or a cross entropy loss function, also named logistic loss function, such aslog Pi + (1 - yt) log(l - pt), where y is a label (ground truth), e.g., 1 for “condition of interest observed” and 0 for “condition of interest not observed” and p is the corresponding estimated probability, e.g., that the input corresponds to the condition of interest.
[0081] After a batch of tensors of any suitable size b, for instance 1 , 50, 256 or a size corresponding to the full size of the training set, has been processed, the parameters of the model can be optimized, for instance, by applying a stochastic gradient descent function if b=1 , a batch gradient descent function if b corresponds to the full size of the training set, or a minibatch gradient descent function in other cases. In some embodiments, one or more gradient descent optimization algorithms such as Adam, NAG or RMSprop can be used. This process is repeated until a configurable amount of time has elapsed, a configurable number of epochs have been run, and / or the model performance measured on the validation dataset has reached a configurable condition.
[0082] In embodiments using a plurality of encoders, for instance convolutional encoders for use with fully connected classifiers, each encoder can be trained individually as a full neural network such as the one illustrated in FIG. 4, then the classifier can be trained on the outputs of the encoders using the appropriate aggregation, e.g., concatenation. It can be appreciated that the classifier shown in FIG. 4 corresponds substantially to a multilayer perceptron taking vector embeddings as input, and it can be trained as such. In embodiments using individual-characterizing data embedding vectors, the encoder(s) can also be trained as a full neural network such as the one illustrated in FIG. 4, then the classifier can be trained on the outputs of the encoders using the appropriate aggregation, e.g., concatenation. In embodiments using an embedding model for individual-characterizing data, the embedding model can be trained on the data using the labels of the dataset, and / or using any suitable similarity measure applicable between individuals, e.g., a similarity measure based in their dataand / or on their known condition(s) or lack thereof. In embodiments using ensemble learning, each model can be trained individually, then the ensemble can be trained. It will be appreciated that other methods of training models are known and can be used with substantially equivalent results.
[0083] With reference again to FIG. 2, in some implementations, the method may include a step of normalizing the chromatograms using a normalizing method compensating for one or more anomalies in the operation of the chromatography system.
[0084] As mentioned above, patterns in the chromatograms can be associated with different pathologies or other conditions as long as the different components of the breath sample are consistently eluting at the same time for each breath sample being analyzed. Different factors, such as anomalies in the chromatography system, can cause a same component in samples processed at different moments by a same GC column to elute with different delays in the chromatography analysis process. In some embodiments, normalizing the chromatograms used for breath analysis, training of the machine learning model or both may improve the reliability of the patterns used by the trained machine learning model. In some cases, a method for anomaly detection and diagnosis in a chromatography system may be used, such as described in international patent application WO2021 / 113977 (GAMACHE et al), the entire contents of which is incorporated herein by reference. Techniques such as described in GAMACHE may be used for monitoring, detecting, diagnosing, and compensating for improper, anomalous, faulty, or suboptimal operation in chromatography systems and methods. The detection and self-diagnosis of anomalies in the operation of a chromatography system may be performed automatically and in real time. In some implementations, such techniques can involve the monitoring and assessment of one or more response factors associated with a chromatography system.
[0085] In one embodiment, a method for anomaly detection and diagnosis in a chromatography system including a sample handling unit, a chromatographicseparation unit, and a detection unit is used. The method can include performing, with the chromatography system, a chromatography analysis of a sample to obtain a chromatogram of the sample, the sample including a known quantity of a reference standard. The method can also include determining peak information corresponding to the reference standard in the chromatogram and determining whether the peak information conforms with an expected response of the chromatography system associated with the reference standard. If the peak information does not conform with the expected response, the method can include determining that there is an anomaly in the operation of the chromatography system and diagnosing a cause of the anomaly as relating to the operation of at least one sample handling unit, the chromatographic separation unit, and the detection unit. Corrective actions can be taken to resolve the anomaly and compensate therefor to ensure reliability and consistency of the position of the peaks appearing on chromatograms acquired with the chromatography system.
[0086] Referring to FIG. 6, in some embodiments a plurality of detectors responding to different molecules or molecule groups are used. The breath sample can be directed to a capture system 610 configured to retain volatile organic compounds. The captured VOCs are then transferred to an injection system 620, which introduces the sample into a gas chromatography column 630. The GC column can separate the VOCs based on their retention times. The eluted compounds are detected by two distinct detectors 640a, b, each generating a 650a, b chromatogram representative of the detected analytes. By way of example, a FID detector may be used for detection of hydrocarbons, a plasma detector may be designed towards specific chemical groups whereas a SCD detector will target sulfur. In some embodiments the signal eluting from a single GC column may be divided such that different portions of this signal are directed to different detectors. An example of two chromatographs obtained from a same GC column using different detectors id show in FIGs. 7A and 7B. The chromatograph 700a of FIG. 7A was obtained using a detector configured to monitor carbon, whereas the chromatograph 700b of FIG. 7B was obtained using a detector configured to monitor the OH chemical group. Both chromatographsmay be processed as images linked to different data sets in the trained machine learning model.
[0087] In other variants, such as shown in FIG. 8, different detectors may receive signals eluting from different GC columns. The breath sample can be directed to a capture system 810 configured to retain volatile organic compounds. The captured VOCs are then transferred to an injection system 820. As biomarkers are not typically well separated by chromatography, in some instances highlighting the presence of a major biomarker may be desired. Different GC columns 830a-c with different properties may provide a different separation of biomarkers. Each column 830a-c can be associated with a detector 840a-c, generating a corresponding chromatograms 850a-c. This may result in the same biomarkers eluting in a different order from separate GC columns, thus multiplying the information while using the same type of detector. Of course, it is possible to use the same principle with detectors of different types. FIG. 9 shows examples of chromatographs of a same sample through two GC columns of different properties. The chromatogram 900a was obtained with a cyanopropyl phenyl-dimethyl polysiloxane based chromatographic column such as the Rxi-624TMchromatographic column, whereas the chromatogram 900b was obtained with a polyethylene glycol based chromatographic column such as the Stabilwax™ chromatographic column. Both types of columns provide proper biomarker separation for quantification, but the biomarker order is different such that they both provide the same information, but in a different view, that is biomarkers, whether known or unknown, will overlap differently and as such, this will provide information in a different way. FIG. 9 illustrates corresponding biomarkers between chromatograms 900a and 900b: in 900a, peak nr. 1 corresponds to pentane, nr. 2 to hexane, nr. 3 to isoprene, nr. 4 to acetone, nr. 5 to 2-propanol, nr. 6 to ethanol, nr. 7 to acetonitrile and nr. 8 to toluene; whereas in 900b, peak nr. 2 corresponds to ethanol, nr. 6 to acetonitrile and nr. 7 to hexane, peraks nrs. 1 , 3-5 and 8 being the same as in 900a.
[0088] In some implementations, generating a plurality of chromatograms for a same breath sample, either through the use of different GC columns, differenttypes of detectors or a combination thereof, improves the robustness of the exhaled breath analysis system. The machine learning model may be trained to accept as input one or more chromatogram associated with a breath sample from an individual, to recognize patterns formed by the peaks in the chromatogram(s), and to infer from the patterns the presence of a condition associated with the individual, for instance, a pathology of the individual.
[0089] In some embodiments, a tridimensional tensor can be generated from the rasterized version(s) of one or more 2D image(s). As an example, matrices generated from a plurality of 2D images as described above can be layered or stacked to create a tridimensional tensor of a suitable shape, e.g., 128 x 128 x 5 corresponding to five 2D images, each rasterized to a size of 128 x 128 pixels. Therefore, two tensor dimensions can be used to represent the pixels of the raster images, and a third tensor dimension can be used to represent each raster image used to generate the tensor.Use Case
[0090] In accordance with an aspect, the methods and systems described herein can be used for the detection of lung cancer using exhaled breath analysis.
[0091] With reference to FIG. 10 (PRIOR ART), a prior art embodiment of a method to perform biomarker analysis using chromatography in a traditional way is shown in accordance with a use case. The biomarkers are captured and contained by a capture device 1010 and then transferred to a chromatographic column 1020 for analysis. The separated biomarkers elute from the column and are detected by a detector 1030, which generates a chromatogram. The detected signal undergoes peak integration and quantification 1040, allowing for the determination of biomarker concentrations. The chromatography is configured in order to minimize biomarker interferences and facilitate their quantification, as shown in FIG. 11 (PRIOR ART). In such an analysis, the biomarkers 1100 were well known and must be properly quantified to be valid and the biomarker levels are used to define if the patient has a lung cancer or not. The difficulty with such amethod is that even if the method is optimized to properly separate the biomarkers 1100, there are still overlaps that are making the quantification inaccurate and consequently the cancer detection, as apparent from FIG. 11 , in which biomarkers A (ethanol) and B (acetone) are in set 1110 of diabetes biomarkers, set 1120 of lung cancer biomarkers set 1130 of heart failure biomarkers and set 1140 of cystic fibrosis biomarkers, biomarkers C (isopropanol) and D (methanol) are in sets 1110, 1120 and 1140, biomarkers E (isoprene), F (propane) and G (undecane) are in sets 1110 and 1120, and only biomarkers H (methul nitrate), I (carbon monoxide), J (toluene), K (m-xylene), L (2,3,4-trimethylhexane), M (2,6,8-trimethylhexane), N (tridecane), O (ethyl benzene), P (2-pentyl nitrate) and Q (ethylene) are in a single set 1110. As an example, FIG. 12 (PRIOR ART) presents a chart showing the lung cancer biomarkers 1210 as well as other biomarkers from other pathologies such as diabetes 1230. The superposition 1220 of both chromatograms 1210 and 1230 shows the overlaps 1225. Additionally, the traditional method is time-consuming. In the use case presented in FIGs. 10 to 12, the required analysis takes in excess of 11 minutes.
[0092] With reference to FIG. 13, an embodiment of a method to perform biomarkers analysis using chromatographic principles as described herein is shown in accordance with the same use case as presented in the preceding paragraphs. The breath sample can be directed to a capture system 1310 configured to retain volatile organic compounds. Two different chromatographic channels 1320a, b may then be used. The channels are configured with columns 1322a, b having different properties in order to provide different biomarker separation, e.g., column 1322a can use Stabilwax™ and column 1322b can use Rxi-624TM. Each channel 1320a, b can be associated with a detector 1324a, b, generating corresponding chromatograms that can be processed by a signal processing algorithm 1330. To use this method, it is not necessary to know the biomarkers for a specific disease. Nonetheless, if some biomarkers are known from the science, the method can be trained with those biomarkers to improve its accuracy. This can enhance the number of reference signals available to train the algorithm. If there are no known biomarkers, the method can be trained withpatients having known conditions, and the resulting signal is a signal containing a combination of shapes associated with the known disease and some shapes related to other biomarkers related to other conditions that are currently unknown. The method may also be configured to speed up the analysis by making the analysis at high chromatographic column temperature.
[0093] In the present use case, the columns were selected following a review of the existing literature that defines potential biomarkers for some specific diseases. A Stabilwax™ and a Rxi-624TMchromatographic columns were selected. The carrier gas is helium and the detectors are Field Enhanced Photo Ionization Detectors (FePid). The columns were selected to offer different biomarker separation properties. Using two chromatographic channels in parallel has the advantage of enhancing the information gathered with respect to having only one chromatographic channel. The table below includes biomarkers association with pathologies or life habits:
[0094] The table below indicates whether the known biomarkers can be quantified without interference or not:
[0095] The tables above are designed to ensure that chromatographic columns having different chromatographic properties are selected.
[0096] In the present use case, using the method described herein, the algorithm is trained with signals that are generated using known lung cancer biomarkers that were identified in the literature. With reference to FIG. 14, the breath sample can be directed to a capture system 1410 configured to retain volatile organic compounds. The sample can then be directed to column 1420a, for instance aStabilwax™ column, associated with first detector 1430a, and column 1420b, for instance a Rxi-624™ column, associated with second detector 1430b. The two generated chromatograms 1435a, b can them be processed by signal processing algorithm 1440. As can be seen in FIG. 14, signals can be used as images, which does not require quantifying the biomarkers. In other words, the biomarkers do not need to be known for the method to be used. As an example, in this use case, thereference gas(es) can include isoprene, methanol, acetone, isopropanol, ethanol and / or undecane.
[0097] In a variation of the present use case illustrated in FIG. 15, the algorithm can also be trained with signals that are generated using known lung cancer biomarkers as well as diabetes biomarkers. The breath sample can be directed to a capture system 1510 configured to retain volatile organic compounds. The sample can then be directed to column 1520a, for instance a Stabilwax™ column, associated with first detector 1530a, and column 1520b, for instance a Rxi-624™ column, associated with second detector 1530b. The two generated chromatograms 1535a, b can them be processed by signal processing algorithm 1540. This offers the benefit of improving the knowledge of the algorithm, considering that the body can be affected by more than one pathology. As an example, in this use case, the reference gas(es) can include methanol, ethanol, isopropanol, acetone, isoprene, toluene, m-xylene, undecane and / or tridecane.
[0098] In yet another variation of the present use case illustrated in FIG. 16, the algorithm can be trained with signals that are generated using know biomarkers associated with a smoking habit. The breath sample can be directed to a capture system 1610 configured to retain volatile organic compounds. The sample can then be directed to column 1620a, for instance a Stabilwax™ column, associated with first detector 1630a, and column 1620b, for instance a Rxi-624™ column, associated with second detector 1630b. The two generated chromatograms 1635a, b can them be processed by signal processing algorithm 1640. This can have the benefit of allowing for the detection of behaviours that could cause interference with the measurement, thereby providing more intelligence to the algorithm. As an example, in this use case, the reference gas(es) can include benzene and / or 2, 5-dimethylfuran.
[0099] In yet another variation of the present use case illustrated in FIG. 17, the algorithm can be trained with samples coming from patients having lung cancer and possibly other unknown pathologies. The breath sample can be directed to acapture system 1710 configured to retain volatile organic compounds. The sample can then be directed to column 1720a, for instance a Stabilwax™ column, associated with first detector 1730a, and column 1720b, for instance a Rxi-624TMcolumn, associated with second detector 1730b. The two generated chromatograms 1735a, b can them be processed by signal processing algorithm 1740. In this use case, the reference gas is the breath.
[0100] In each use case of the method presented herein, a CNN 1830 can be trained with reference signals as images 1810, as shown in FIG. 18. Other relevant information 1820 can additionally be used to train the CNN 1830, including for instance the CO2 level in the exhaled breath, the total VOC level in the exhaled breath, and / or the age, gender, weight and / or known pathologies of a training-time subject.
[0101] The CNN thus trained can be used to perform inferences about an inference-time subject in order to detect cancer in the subject, as shown in FIG. 19. This can include receiving real signals 1910 from a patient potentially having lung cancer, and providing them as input to the CNN 1930, which is then able to compute and output a vector of inferences 1940 regarding pathologies identified in the real signals. Other relevant information 1920 can additionally be used to perform inference in the CNN, including for instance the CO2 level in the exhaled breath, the total VOC level in the exhaled breath, and / or the age, gender, weight and / or known pathologies of an inference-time subject.
[0102] In summary, the method described herein differentiates itself from traditional chromatography from the following points. There is no need to know the biomarkers associated with a pathology as the algorithm described herein does not quantify, but rather looks for signal variations at specific location in time (elution time). The method relies on the fact that for known chromatographic conditions, biomarkers will always generate a signal variation at a specific point in time. In other words, the algorithm is looking for signal variations instead of quantifying biomarkers, using a method similar to pattern identification in 2D images. Thebiomarkers do not need to be properly separated for quantification as normally required in chromatography. FIG. 20A and 20B show the traditional and the novel methods side by side. In the prior art method shown in FIG. 20A, a reference mixture with known biomarkers is prepared 2010, and the elution times of these biomarkers are identified on a chromatographic column 2020. The instrument is then calibrated using the known biomarkers 2030. A sample from a cancerous patient is injected into the system 2040, and the biomarkers in the sample, e.g., benzene at 23 ppb and isoprene at 95 ppb, are quantified 2080a. The quantified biomarker levels are compared with predefined threshold levels 2085, and the presence of cancer is identified 2090 based on this comparison. In the improved method shown in FIG. 20B, a model is trained with known biomarkers 2050 if they are known, and then trained with patients having known pathologies 2060. A sample from a cancerous patient is injected into the system 2070, and pattern analysis is performed in the generated chromatograms 2080b, and the presence of cancer 2090 is identified based on this pattern analysis.
[0103] In some implementations, the power of using multiple parallel chromatographic channels having different characteristics (detector, column) can be used to increase the level of information obtained by an analysis, allowing for a multi-dimensional analysis. For example, a certain biomarker may interfere with other biomarkers on the first chromatographic channel, but be well separated in the other channel. In such a case, the biomarker of interest that is superimposed by another unknown biomarker will generate two signals, one on each channel. One signal will have interference but still a variation while the other channel will have a signal associated do that specific pathologies. FIGs. 21 A and 22B show an example illustrating the power of using multiple parallel chromatographic channels to enhance the data. In this case, acetone is only well separated in signal 2 (2120b in FIG. 21 B). In signal 1 (2120a in FIG. 21 A), it is merged with other biomarkers. Both methods, however, offer a clean signal 2110 for isoprene. Using a traditional chromatographic approach, it would be necessary to either slow down the analysis, or to use a chromatographic column that separates well acetone from the other biomarkers. However, this new column may cause other biomarkers to merge,which is why multiple channels are important. Both signals will have an acetone signal, but only signal 2 will have one that is specific to acetone. The variation of both signals will be taken into consideration by the learning algorithm. When the two signals are combined together in the algorithm, they enhanced the precision due to the enhanced data. The present method is also much faster, as it does not require quantification, and instead only identifies signal variation associated with the biomarkers.
[0104] Of course, numerous additional modifications could be made to the embodiments described above without departing from the scope of protection as defined in the appended claims.
Claims
CLAIMS1 . A method for analysis of a breath sample, comprising: using a chromatography system comprising a chromatography column and one or more detectors, obtaining at least one chromatogram representative of an elution of components of the breath sample from the chromatography column over time, each of the at least one chromatogram defining a 2D image of a plurality of peaks as a function of time; and running a trained machine learning model at least on the at least one chromatogram to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one chromatogram, wherein the trained machine learning model has been trained with a dataset comprising a plurality of training pairs, each of the plurality of training pairs comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
2. The method of claim 1 , wherein the at least one chromatogram is a plurality of chromatograms obtained via a plurality of detectors.
3. The method of claim 1 or 2, wherein the machine learning model comprises: convolutional layers defining an encoder, a first convolutional layer being configured to accept as input the at least one chromatogram, and a last convolutional layer being configured to generate as output at least one vector embedding representing the plurality of peaks; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least the at least one vector embedding, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample.
4. The method of claim 3, wherein the encoder is configured to accept as input a tridimensional tensor, the method further comprising generating the tridimensional tensor by layering images defined by the at least one chromatogram.
5. The method of claim 1 or 2, wherein the machine learning model comprises: convolutional layers defining a plurality of encoders, each encoder being configured to accept as input exactly one chromatogram of the at least one chromatogram and to generate as output a corresponding exactly one vector embedding representing a portion of the plurality of peaks associated with the one of the one or more detectors, the plurality of encoders thereby generating a plurality of vector embeddings; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least an aggregation of the plurality of vector embeddings, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample, the method further comprising aggregating the plurality of vector embeddings.
6. The method of any one of claims 3 to 5, further comprising obtaining data characterizing an individual associated with the breath sample, wherein each of the plurality of training pairs further comprises training data, and wherein the machine learning model is run further on the data characterizing the individual.
7. The method of claim 6, further comprising, using a trained embedding model, generating a data vector embedding based on the data characterizing the individual, wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector embedding.
8. The method of claim 6, further comprising:converting each datum of the data characterizing the individual in a predefined numeric format; and combining the converted data characterizing the individual in a data vector, wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector.
9. A method for analysis of a breath sample, comprising: using an analytical system comprising one or more detectors, obtaining at least one graph representing a distribution of components of the breath sample as a 2D image of a plurality of peaks; and running a trained machine learning model at least on the at least one graph to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one graph, wherein the trained machine learning model has been trained with a dataset comprising training pairs each comprising at least one 2D training image obtained from a training breath sample and at least one label corresponding to a training pathology.
10. The method of claim 9, wherein the at least one graph is a plurality of graphs obtained via a plurality of detectors.11 . The method of claim 9 or 10, wherein the analytical system is a chromatography system comprising a chromatography column, wherein the at least one graph corresponds to at least one chromatogram representative of an elution of components of the breath sample from the chromatography column over time, and wherein each of the at least one 2D image is an image of the plurality of peaks as a function of time.
12. The method of any one of claims 9 to 11 , further comprising obtaining data characterizing an individual associated with the breath sample, wherein eachof the plurality of training pairs further comprises training data, and wherein the machine learning model is run further on the data characterizing the individual.
13. The method of claim 12, wherein the data characterizing the individual comprises at least one of an age, a weight, a height, a current vital sign, a historical vital sign, a current symptom, a historical symptom, body fluid analysis results, and known medical conditions.
14. The method of claim 12 or 13, further comprising generating at least one raster image associated with the at least one 2D image, each of the at least one raster image corresponding to a matrix comprising a configurable number of matrix cells, such that each respective cell of the matrix cells comprises a value indicative of whether a line is visible in a pixel corresponding to the respective cell.
15. The method of claim 14, wherein the machine learning model is configured to accept as input at least the at least one raster image.
16. The method of claim 15, wherein the machine learning model comprises: convolutional layers defining an encoder, a first convolutional layer being configured to accept as input at least one of the at least one raster image, and a last convolutional layer being configured to generate as output at least one vector embedding representing the plurality of peaks; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least the at least one vector embedding, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample.
17. The method of claim 16, wherein the encoder is configured to accept as input a tridimensional tensor, the method further comprising generating the tridimensional tensor by layering the at least one raster image.
18. The method of claim 15, wherein the machine learning model comprises: convolutional layers defining a plurality of encoders, a first convolutional layer being configured to accept as input exactly one raster image of the at least one raster image associated with one of the one or more detectors, and a last convolutional layer being configured to generate as output a corresponding exactly one vector embedding representing a portion of the plurality of peaks associated with the one of the one or more detectors, the plurality of encoders thereby generating a plurality of vector embeddings; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least an aggregation of the plurality of vector embeddings, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample, the method further comprising aggregating the plurality of vector embeddings.
19. The method of claim 18, wherein the plurality of encoders are identical.
20. The method of any one of claims 16 to 19, further comprising, using a trained embedding model, generating a data vector embedding based on the data characterizing the individual, wherein the first connected layer of the at least one connected layer being configured to accept as additional input the data vector embedding.21 . The method of any one of claims 16 to 19, further comprising: converting each datum of the data characterizing the individual in a predefined numeric format; and combining the converted data characterizing the individual in a data vector,wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector.
22. The method of any one of claims 9 to 21 , wherein the at least one probability comprises a first probability associated with the presence of the pathology and a second probability associated with an absence of the pathology.
23. The method of any one of claims 9 to 21 , wherein the at least one probability comprises at least one additional probability associated with the presence and / or absence of at least one additional pathology.
24. The method of any one of claims 9 to 23, wherein at least some of the at least one training graph are associated with training breath samples characterizing a smoking habit.
25. The method of any one of claims 9 to 24, wherein at least some of the at least one training graph are associated with training breath samples characterizing lung cancer and / or diabetes.
26. The method of any one of claims 1 to 25, wherein the pathology is one of cancer, respiratory disease, pulmonary disease, kidney disease, liver diseases, diabete, alcohol intoxication, organ rejection, sleep apnea, and mental and / or physical stress.
27. A system for analysis of a breath sample, comprising: a chromatography system comprising a chromatography column and one or more detectors, configured to obtain at least one chromatogram representative of an elution of components of the breath sample from the chromatography column over time, each of the at least one chromatogram defining a 2D image of a plurality of peaks as a function of time; and a trained machine learning model configured to run at least on the at least one chromatogram to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one chromatogram,wherein the trained machine learning model has been trained with a dataset comprising a plurality of training pairs, each of the plurality of training pairs comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
28. The system of claim 27, wherein the one or more detectors is a a plurality of detectors configured to obtain a plurality of chromatograms.
29. The system of claim 27 or 28, wherein the machine learning model comprises: convolutional layers defining an encoder, a first convolutional layer being configured to accept as input the at least one chromatogram, and a last convolutional layer being configured to generate as output at least one vector embedding representing the plurality of peaks; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least the at least one vector embedding, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample.
30. The system of claim 29, wherein the encoder is configured to accept as input a tridimensional tensor, further comprising generating the tridimensional tensor by layering images defined by the at least one chromatogram.
31. The system of claim 27 or 28, wherein the machine learning model comprises: convolutional layers defining a plurality of encoders, each encoder being configured to accept as input exactly one chromatogram of the at least one chromatogram and to generate as output a corresponding exactly one vector embedding representing a portion of the plurality of peaks associated with the one of the one or more detectors, the plurality of encoders thereby generating a plurality of vector embeddings; andat least one connected layer defining a classifier, a first connected layer being configured to accept as input at least an aggregation of the plurality of vector embeddings, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample.
32. The system of any one of claims 29 to 31 , further comprising a records system configured to provide data characterizing an individual associated with the breath sample, wherein each of the plurality of training pairs further comprises training data, and wherein the machine learning model is run further on the data characterizing the individual.
33. The system of claim 32, further comprising a trained embedding model configured to generate a data vector embedding based on the data characterizing the individual, wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector embedding.
34. The system of claim 33, further configured to: convert each datum of the data characterizing the individual in a predefined numeric format; and combine the converted data characterizing the individual in a data vector, wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector.
35. A system for analysis of a breath sample, comprising: an analytical system comprising one or more detectors configured to obtain at least one graph representing a distribution of components of the breath sample as a 2D image of a plurality of peaks; anda trained machine learning model configured to run at least on the at least one graph to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one graph, wherein the trained machine learning model has been trained with a dataset comprising a training pairs, each comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
36. The system of claim 35, wherein the at least one graph is a plurality of graphs obtained via a plurality of detectors.
37. The system of claim 35 or 36, wherein the analytical system is a chromatography system comprising a chromatography column, wherein the at least one graph corresponds to at least one chromatogram representative of an elution of components of the breath sample from the chromatography column over time, and wherein each of the at least one 2D image is an image of the plurality of peaks as a function of time.
38. The system of any one of claims 35 to 37, further comprising a records system configured to provide data characterizing an individual associated with the breath sample, wherein each of the plurality of training pairs further comprises training data, and wherein the machine learning model is run further on the data characterizing the individual.
39. The system of claim 38, wherein the data characterizing the individual comprises at least one of an age, a weight, a height, a current vital sign, a historical vital sign, a current symptom, a historical symptom, body fluid analysis results, and known medical conditions.
40. The system of claim 38 or 39, further comprising a rasterization module configured to generate at least one raster image associated with the at least one 2D image, each of the at least one raster image corresponding to a matrix comprising a configurable number of matrix cells, such that each respective cellof the matrix cells comprises a value indicative of whether a line is visible in a pixel corresponding to the respective cell.
41. The system of claim 40, wherein the machine learning model is configured to accept as input at least the at least one raster image.
42. The system of claim 41 , wherein the machine learning model comprises: convolutional layers defining an encoder, a first convolutional layer being configured to accept as input at least one of the at least one raster image, and a last convolutional layer being configured to generate as output at least one vector embedding representing the plurality of peaks; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least the at least one vector embedding, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample.
43. The system of claim 42, wherein the encoder is configured to accept as input a tridimensional tensor, further comprising generating the tridimensional tensor by layering the at least one raster image.
44. The system of claim 41 , wherein the machine learning model comprises: convolutional layers defining a plurality of encoders, a first convolutional layer being configured to accept as input exactly one raster image of the at least one raster image associated with one of the one or more detectors, and a last convolutional layer being configured to generate as output a corresponding exactly one vector embedding representing a portion of the plurality of peaks associated with the one of the one or more detectors, the plurality of encoders thereby generating a plurality of vector embeddings; andat least one connected layer defining a classifier, a first connected layer being configured to accept as input at least an aggregation of the plurality of vector embeddings, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample.
45. The system of claim 44, wherein the plurality of encoders are identical.
46. The system of any one of claims 42 to 45, further comprising a trained embedding model configured to generate a data vector embedding based on the data characterizing the individual, wherein the first connected layer of the at least one connected layer being configured to accept as additional input the data vector embedding.
47. The system of any one of claims 42 to 45, further configured to: convert each datum of the data characterizing the individual in a predefined numeric format; and combine the converted data characterizing the individual in a data vector, wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector.
48. The system of any one of claims 35 to 47, wherein the at least one probability comprises a first probability associated with the presence of the pathology and a second probability associated with an absence of the pathology.
49. The system of any one of claims 35 to 48, wherein the at least one probability comprises at least one additional probability associated with the presence and / or absence of at least one additional pathology.
50. The system of any one of claims 35 to 49, wherein at least some of the at least one training graph are associated with training breath samples characterizing a smoking habit.51 . The system of any one of claims 35 to 50, wherein at least some of the at least one training graph are associated with training breath samples characterizing lung cancer and / or diabetes.
52. The system of any one of claims 27 to 51 , wherein the pathology is one of cancer, respiratory disease, pulmonary disease, kidney disease, liver diseases, diabete, alcohol intoxication, organ rejection, sleep apnea, and mental and / or physical stress.
53. A non-transitory computer-readable medium having instructions stored thereon which, when executed by one or more processors, cause the one or more processors to: using a chromatography system comprising a chromatography column and one or more detectors, obtain at least one chromatogram representative of an elution of components of a breath sample from the chromatography column over time, each of the at least one chromatogram defining a 2D image of a plurality of peaks as a function of time; and run a trained machine learning model at least on the at least one chromatogram to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one chromatogram, wherein the trained machine learning model has been trained with a dataset comprising a plurality of training pairs, each of the plurality of training pairs comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
54. The computer-readable medium of claim 53, wherein the at least one chromatogram is a plurality of chromatograms obtained via a plurality of detectors.
55. The computer-readable medium of claim 53 or 54, wherein the machine learning model comprises:convolutional layers defining an encoder, a first convolutional layer being configured to accept as input the at least one chromatogram, and a last convolutional layer being configured to generate as output at least one vector embedding representing the plurality of peaks; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least the at least one vector embedding, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample.
56. The computer-readable medium of claim 55, wherein the encoder is configured to accept as input a tridimensional tensor, further comprising generating the tridimensional tensor by layering images defined by the at least one chromatogram.
57. The computer-readable medium of claim 53 or 54, wherein the machine learning model comprises: convolutional layers defining a plurality of encoders, each encoder being configured to accept as input exactly one chromatogram of the at least one chromatogram and to generate as output a corresponding exactly one vector embedding representing a portion of the plurality of peaks associated with the one of the one or more detectors, the plurality of encoders thereby generating a plurality of vector embeddings; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least an aggregation of the plurality of vector embeddings, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample, wherein the instructions further cause the one or more processors to aggregate the plurality of vector embeddings.
58. The computer-readable medium of any one of claims 55 to 57, wherein the instructions further cause the one or more processors to obtain data characterizing an individual associated with the breath sample, wherein each of the plurality of training pairs further comprises training data, and wherein the machine learning model is run further on the data characterizing the individual.
59. The computer-readable medium of claim 58, wherein the instructions further cause the one or more processors to, using a trained embedding model, generate a data vector embedding based on the data characterizing the individual, wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector embedding.
60. The computer-readable medium of claim 59, wherein the instructions further cause the one or more processors to: convert each datum of the data characterizing the individual in a predefined numeric format; and combine the converted data characterizing the individual in a data vector, wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector.61 . A non-transitory computer-readable medium having instructions stored thereon which, when executed by one or more processors, cause the one or more processors to: using an analytical system comprising one or more detectors, obtain at least one graph representing a distribution of components of a breath sample as a 2D image of a plurality of peaks; and run a trained machine learning model at least on the at least one graph to infer a presence of a pathology in the breath sample based on patterns formed by the peaks in the at least one graph,wherein the trained machine learning model has been trained with a dataset comprising a training pairs, each comprising at least one 2D training image obtained from a training breath sample, and at least one label corresponding to a training pathology.
62. The computer-readable medium of claim 61 , wherein the at least one graph is a plurality of graphs obtained via a plurality of detectors.
63. The computer-readable medium of claim 61 or 62, wherein the analytical system is a chromatography system comprising a chromatography column, wherein the at least one graph corresponds to at least one chromatogram representative of an elution of components of the breath sample from the chromatography column over time, and wherein each of the at least one 2D image is an image of the plurality of peaks as a function of time.
64. The computer-readable medium of any one of claims 61 to 63, wherein the instructions further cause the one or more processors to obtain data characterizing an individual associated with the breath sample, wherein each of the plurality of training pairs further comprises training data, and wherein the machine learning model is run further on the data characterizing the individual.
65. The computer-readable medium of claim 64, wherein the data characterizing the individual comprises at least one of an age, a weight, a height, a current vital sign, a historical vital sign, a current symptom, a historical symptom, body fluid analysis results, and known medical conditions.
66. The computer-readable medium of claim 64 or 65, wherein the instructions further cause the one or more processors to generate at least one raster image associated with the at least one 2D image, each of the at least one raster image corresponding to a matrix comprising a configurable number of matrix cells, such that each respective cell of the matrix cells comprises a value indicative of whether a line is visible in a pixel corresponding to the respective cell.
67. The computer-readable medium of claim 66, wherein the machine learning model is configured to accept as input at least the at least one raster image.
68. The computer-readable medium of claim 67, wherein the machine learning model comprises: convolutional layers defining an encoder, a first convolutional layer being configured to accept as input at least one of the at least one raster image, and a last convolutional layer being configured to generate as output at least one vector embedding representing the plurality of peaks; and at least one connected layer defining a classifier, a first connected layer being configured to accept as input at least the at least one vector embedding, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample.
69. The computer-readable medium of claim 68, wherein the encoder is configured to accept as input a tridimensional tensor, further comprising generating the tridimensional tensor by layering the at least one raster image.
70. The computer-readable medium of claim 67, wherein the machine learning model comprises: convolutional layers defining a plurality of encoders, a first convolutional layer being configured to accept as input exactly one raster image of the at least one raster image associated with one of the one or more detectors, and a last convolutional layer being configured to generate as output a corresponding exactly one vector embedding representing a portion of the plurality of peaks associated with the one of the one or more detectors, the plurality of encoders thereby generating a plurality of vector embeddings; andat least one connected layer defining a classifier, a first connected layer being configured to accept as input at least an aggregation of the plurality of vector embeddings, and a last connected layer being configured to generate as output at least one probability associated with the presence of the pathology in the breath sample, wherein the instructions further cause the one or more processors to aggregate the plurality of vector embeddings.
71. The computer-readable medium of claim 70, wherein the plurality of encoders are identical.
72. The computer-readable medium of any one of claims 68 to 71 , wherein the instructions further cause the one or more processors to, using a trained embedding model, generate a data vector embedding based on the data characterizing the individual, wherein the first connected layer of the at least one connected layer being configured to accept as additional input the data vector embedding.
73. The computer-readable medium of any one of claims 68 to 71 , wherein the instructions further cause the one or more processors to: convert each datum of the data characterizing the individual in a predefined numeric format; and combine the converted data characterizing the individual in a data vector, wherein the first connected layer of the at least one connected layer is configured to accept as additional input the data vector.
74. The computer-readable medium of any one of claims 61 to 73, wherein the at least one probability comprises a first probability associated with the presence of the pathology and a second probability associated with an absence of the pathology.
75. The computer-readable medium of any one of claims 61 to 74, wherein the at least one probability comprises at least one additional probability associated with the presence and / or absence of at least one additional pathology.
76. The computer-readable medium of any one of claims 61 to 75, wherein at least some of the at least one training graph are associated with training breath samples characterizing a smoking habit.
77. The method of any one of claims 61 to 76, wherein at least some of the at least one training graph are associated with training breath samples characterizing lung cancer and / or diabetes.
78. The method of any one of claims 53 to 77, wherein the pathology is one of cancer, respiratory disease, pulmonary disease, kidney disease, liver diseases, diabete, alcohol intoxication, organ rejection, sleep apnea, and mental and / or physical stress.