Method and apparatus for classifying carbon dioxide oscillogram and predicting cardiopulmonary disease using same
By extracting turning points in carbon dioxide waveforms and applying machine learning models, the classification problem of carbon dioxide measurement in the diagnosis of cardiopulmonary diseases was solved, the automated analysis and accurate prediction of carbon dioxide waveforms were achieved, and the diagnostic efficiency and accuracy of cardiopulmonary diseases were improved.
Patent Information
- Application Number
- CN202380089037.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2023-11-07
- Publication Date
- 2025-09-12
AI Technical Summary
Existing carbon dioxide measurement methods lack a rapid and effective classification method in the diagnosis of cardiopulmonary diseases, and it is difficult to accurately distinguish between healthy and diseased carbon dioxide waveforms, especially for early or asymptomatic COPD patients.
By extracting the turning points of the respiratory waveform (δ, γ, α turning points) from the carbon dioxide waveform and classifying these features using a trained machine learning model, an automated analysis method for the carbon dioxide waveform is established to predict the presence, type, and severity of cardiopulmonary diseases.
It achieves accurate classification and automated analysis of carbon dioxide waveforms, improves the efficiency and accuracy of diagnosis of cardiopulmonary diseases, can quickly identify health status or disease severity, and is suitable for local or remote computing environments.
Smart Images

Figure CN120641040A_ABST
Abstract
Description
Background Art
[0001] Cardiopulmonary disease, which refers to cardiovascular disease and respiratory diseases, is the leading cause of premature death worldwide. Cardiovascular disease refers to conditions that affect the heart or blood circulation. Respiratory disease, also known as lung disease, refers to a variety of conditions that make breathing difficult for those with the condition. Asthma and chronic obstructive pulmonary disease (COPD) are two common examples, with COPD referring to a group of obstructive respiratory diseases that primarily include emphysema and chronic bronchitis. COPD is a progressive disease with no known cure and has been reported to be the third most common cause of death worldwide. This has led to COPD being widely studied worldwide as part of efforts to improve early diagnosis and treatment to reduce its impact on patients' quality of life.
[0002] The current standard for COPD diagnosis is spirometry, which relies on the patient's ability to forcefully exhale. However, the main disadvantage of this measurement is that it relies on the subject's effort and is therefore unreliable and non-specific, particularly for COPD patients who may be asymptomatic or have early-stage disease. Capnometry is an alternative technique that can be used to assess lung health by measuring the partial pressure of carbon dioxide in inspired and expired air. An advantage of capnometry over spirometry is that it can be used with tidal breathing and therefore only requires the patient to breathe naturally. Capnometry produces a tidal breathing record (also known as a capnogram) with a series of respiratory waveforms, and the characteristics of these waveforms can differ significantly between healthy patients and those with cardiopulmonary diseases (e.g., COPD, asthma, and heart failure), depending on the severity of their condition.
[0003] The use of capnography to predict cardiopulmonary disease is still in its infancy, and therefore requires continued development in many areas before it can be widely adopted. For example, which features and characteristics of a capnogram are most important for classifying capnograms and predicting cardiopulmonary disease? How can these characteristics be determined to help predict the presence of cardiopulmonary disease in a patient? And how can this be done quickly and efficiently to provide a method that is feasible in the real world? The present invention provides solutions to these problems. SUMMARY OF THE INVENTION
[0005] According to a first aspect, the present disclosure provides a method for classifying a carbon dioxide waveform graph, the method comprising: obtaining a respiratory waveform from a carbon dioxide waveform graph; determining one or more turning points of the respiratory waveform, wherein the turning points comprise: a delta turning point located between an expiratory baseline and an expiratory ascending limb, a gamma turning point located between an inspiratory descending limb and an inspiratory baseline, and an alpha turning point located between an expiratory ascending limb and an expiratory plateau; extracting features of the respiratory waveform using the turning points; and applying a trained machine learning model to the extracted features using the extracted features, wherein the trained machine learning model is configured to output a classification of the carbon dioxide waveform graph based on the features.
[0006] Capnography classification can be used in a variety of ways. For example, by highlighting and / or classifying capnography for further analysis, or ranking capnography according to risk or important factors associated with the capnography. This can include predicting cardiopulmonary conditions or health status to rule out such conditions, predicting the severity or subtype of a condition, predicting medication response, or predicting other clinical outcomes, such as the probability of hospitalization within a given time window. Using a trained machine learning model in this way provides accurate and automated capnography classification, which has not been previously achieved.
[0007] In some embodiments of the first aspect, classification of capnography can be used to predict cardiopulmonary disease. These can be considered methods for predicting whether a capnography waveform is generated by a user with cardiopulmonary disease, wherein the trained machine learning model is configured to predict cardiopulmonary disease.
[0008] That is, an embodiment of the first aspect includes a method for predicting whether a carbon dioxide waveform graph is generated by a user suffering from cardiopulmonary disease, the method comprising: obtaining a respiratory waveform from the carbon dioxide waveform graph; determining one or more turning points of the respiratory waveform, wherein the turning points include: a δ turning point, which is located between the expiratory baseline and the expiratory ascending limb; a γ turning point, which is located between the inspiratory descending limb and the inspiratory baseline; and an α turning point, which is located between the expiratory ascending limb and the expiratory platform, using the turning points to extract features of the respiratory waveform; and using the extracted features, applying a trained machine learning model to the extracted features, wherein the trained machine learning model is configured to predict cardiopulmonary disease based on the features.
[0009] In particular, the method can be considered for predicting cardiopulmonary disease based on or using a capnography waveform, that is, whether the capnography waveform is generated by a diseased respiratory tract and / or circulatory system (herein referred to as the circulatory system). One objective is to use data generated from a device configured to measure CO2 concentration to develop a classifier that distinguishes between capnography waveforms generated by healthy lungs and those generated by lungs or airways with cardiopulmonary disease. For example, a cardiopulmonary disease can be COPD, pulmonary fibrosis, asthma, asthma-COPD overlap syndrome (ACOS), or small airways disease. Throughout this application, we may refer to diseases. While we also refer to subtypes of the same disease, for the sake of brevity, only the disease will be mentioned. Predicting cardiopulmonary disease can be considered as predicting a state of health and therefore includes predicting the absence of cardiopulmonary disease, i.e., referring to a healthy patient or a patient without cardiopulmonary disease (but suffering from another type of disease). Predicting cardiopulmonary disease based on the features also includes predicting the type of cardiopulmonary disease, subtype of cardiopulmonary disease, and / or severity of the cardiopulmonary disease based on the features. This information can be represented by classification of the capnography waveform.
[0010] The prediction method and the training method using the trained model can each be performed at a local computer or a remote computer, remote from the patient associated with the capnogram. The steps of the method can be performed entirely by a local processing unit, a remote processing unit (e.g., in the cloud), or split between the local and remote processing units.
[0011] The respiratory waveform may be a respiratory waveform or a representation thereof.
[0012] In this example, the waveform is a digitally sampled signal representing a quantized amplitude signal. This signal can be encoded as an array of floating-point samples, each representing the amplitude of the respiratory signal at a point in time. The samples may not directly correspond to the amplitude of the respiratory waveform, but should be understood as a digital representation of the respiratory waveform. Similarly, the waveform can be encoded in any suitable manner, such as an array of binary vectors or matrices.
[0013] A turning point is a sample of digital samples corresponding to turning points in the waveform. This method can identify a single sample or element in the code, the index of a sample, or the time point corresponding to the sample. Each sample can generally correspond to a CO2 value at that point in time.
[0014] The trained machine learning model is trained using the same or corresponding extracted features as the extracted features, which have been extracted from turning points of other carbon dioxide waveforms that are marked as positively or negatively correlated with cardiopulmonary diseases and in some instances with the severity or subtype of the diseases, or extracted based on turning points of the carbon dioxide waveforms. For example, the trained machine learning model can be trained using a method according to any embodiment of the second aspect below. As used herein, the term machine learning is also used to refer to statistical inference. For example, a statistical inference model can be referred to as a machine learning model, and the step of training a machine learning model can refer to training a statistical inference model (or a machine learning model).
[0015] Determining the turning points of the waveform enables accurate and consistent feature extraction across different segments of the respiratory waveform, which in turn enables a trained machine learning model to more accurately and interpretably predict whether a capnography waveform including a respiratory waveform was generated by a user with a cardiopulmonary disease. The delta, gamma, and alpha turning points of the respiratory waveform have been found to be particularly beneficial for capnography prediction accuracy.
[0016] The turning point of the waveform is a parameter that is relevant to the prediction of multiple different cardiopulmonary diseases, such as COPD and asthma. Therefore, when predicting different cardiopulmonary diseases based on a single respiratory waveform, these steps do not need to be repeated, further improving the efficiency of the method, especially at a larger scale.
[0017] Capnography and respiratory waveforms generated by different airways can vary significantly, and can also vary significantly from the same airway at different times or under different conditions (e.g., when affected by cardiopulmonary disease). These variations have created difficulties when attempting to accurately analyze capnography quickly and / or efficiently. Automated capnography analysis is particularly challenging due to the vast degree of variability between respiratory waveforms and the various factors that contribute to these variations (e.g., cardiopulmonary disease, age, weight, time of day, medication use, smoking status, location, presence of other medical conditions, etc.). However, identifying turning points in the waveforms and using these turning points as the basis for feature extraction and application of machine learning models can be easily automated in a consistent manner to improve prediction efficiency while still maintaining consistently high prediction accuracy.
[0018] When turning points mark the transition between different phases of the respiratory waveform (e.g., the alpha turning point is between the ascending limb of exhalation and the plateau of exhalation), identifying any of these points also improves the interpretability of the method. That is, in addition to predicting cardiopulmonary disease based on the extracted features, the trained machine learning model can also indicate which phase of the respiratory waveform contributes to the prediction and the degree to which the phase contributes to the prediction.
[0019] The terms capnography and respiration recording are used interchangeably herein to refer to continuous capnography measurements during a single monitoring session. For example, a capnography waveform ends when a user or operator stops recording the user's breathing with a capnography monitor, and begins recording again when another capnography waveform begins. However, capnography and respiration recording can also refer to a given period of continuous CO2 recording (e.g., in a ventilator circuit).
[0020] Skilled artisans will appreciate that the inspiratory baseline and the expiratory baseline can be considered the same baseline of a capnogram / respiratory recording, i.e., the respiratory baseline or capnogram baseline. The inspiratory baseline and the expiratory baseline are named differently to distinguish between the baseline adjacent to the expiratory time period of a respiratory waveform and the baseline adjacent to the inspiratory time period of a respiratory waveform (when the individual respiratory waveforms have been separated), and to more clearly delineate the turning points of the waveforms.
[0021] According to a second aspect, a method for training a machine learning model to learn indicators of a capnography waveform is provided, wherein for a plurality of capnography waveforms, the method comprises: i) obtaining a respiratory waveform from each capnogram; ii) determining one or more turning points of the respiratory waveform, wherein the turning points include: a delta turning point between the expiratory baseline and the expiratory ascending limb, a gamma turning point between the inspiratory descending limb and the inspiratory baseline, and an alpha turning point between the expiratory ascending limb and the expiratory plateau; iii) extracting features of respiratory waveform using multiple turning points; A label for each capnogram is obtained to indicate a classification of the capnogram; and a machine learning model is trained using the extracted features of the plurality of respiratory waveforms and the corresponding labels of the corresponding capnograms to learn indicators of the capnograms to create a classification function.
[0022] The plurality of capnograms may be individual capnograms from a plurality of different users and / or a plurality of capnograms generated by a user at different times. The plurality of respiratory waveforms may also be obtained from a single capnogram in the plurality of capnograms.
[0023] Step iii) of the second aspect generates a characterized respiratory waveform or capnography. Steps i) to iii) may be repeated to obtain extracted features of a plurality of respiratory waveforms (ie, to generate a plurality of characterized respiratory waveforms).
[0024] In some embodiments of the second aspect, a machine learning model is trained to learn indicators of cardiopulmonary disease (i.e., indicators of capnograms generated from users with cardiopulmonary disease) to create a prediction function, wherein the capnogram classification indicates whether the capnogram is from a diseased respiratory tract and / or circulatory system (i.e., each label is used to indicate whether the capnogram is from a diseased respiratory tract and / or circulatory system), and the training module creates the prediction function.
[0025] That is, the embodiment of the second aspect includes a method for training a machine learning model to learn indicators of a capnogram generated by a user with a cardiopulmonary disease, the method comprising, for a plurality of capnograms,: i) obtaining a respiratory waveform from each capnogram; ii) determining one or more turning points of the respiratory waveform, wherein the turning points include: a delta turning point between the expiratory baseline and the expiratory ascending limb, a gamma turning point between the inspiratory descending limb and the inspiratory baseline, and an alpha turning point between the expiratory ascending limb and the expiratory plateau; iii) extracting features of respiratory waveform using multiple turning points; Obtain a label for each capnogram, the label indicating whether the capnogram is from a diseased respiratory tract and / or circulatory system; and, using the extracted features of the plurality of respiratory waveforms and the corresponding labels of the corresponding capnograms, train a machine learning model to learn indicators of cardiopulmonary disease to create a predictive function.
[0026] The machine learning model can be trained to learn indicators of carbon dioxide waveforms corresponding to several different cardiopulmonary diseases. Optionally, the label of each carbon dioxide waveform should indicate whether the carbon dioxide waveform is from the respiratory tract affected by the cardiopulmonary disease currently being trained, and can also indicate the subtype of the cardiopulmonary disease, the severity of the disease, drug responsiveness and / or other clinical outcomes. For example, during the training of the COPD classification model, the label of each carbon dioxide waveform can indicate whether the carbon dioxide waveform is from the respiratory tract with COPD lesions or the respiratory tract with non-COPD lesions. In the example of training the asthma classification model, the label of each carbon dioxide waveform can indicate whether the carbon dioxide waveform is from the asthmatic respiratory tract or the non-asthmatic respiratory tract. The machine learning model can also be trained to learn indicators of carbon dioxide waveforms produced by healthy users who do not suffer from any cardiopulmonary disease. Preferably, the model is trained using carbon dioxide waveforms corresponding to healthy and diseased users to provide more general and accurate functions.
[0027] The first and second aspects may include any one of the following embodiments: Preferably, the respiratory waveform represents a single respiratory cycle. Obtaining the respiratory waveform may include segmenting the capnogram into a plurality of capnogram portions, wherein each capnogram portion represents a single respiratory waveform corresponding to a single respiratory cycle.
[0028] A single respiratory cycle should be a completely recorded breath. That is, a breath is recorded from exhalation to inhalation or from inhalation to exhalation, with both exhalation and inhalation fully recorded. A respiratory waveform / respiratory cycle can also correspond to a partial breath (e.g., a recording of capnography data that starts or stops mid-breath). Even when a respiratory waveform corresponds to a partial breath, turning points can still be determined using the techniques of the present invention (and those determined points can be used to extract features). However, depending on when the recording of the capnography data is started or stopped, as many turning points may not be determined or as many features extracted as for a respiratory waveform corresponding to a full breath. Therefore, respiratory waveforms corresponding to partial breaths can be identified and calculated (e.g., filtered out from analysis). Having a respiratory waveform represent a single respiratory cycle reduces the amount of computation required to train a machine learning model to accurately predict cardiopulmonary disease, while still providing the required information (e.g., turning points). This also facilitates further analysis, such as average waveform generation and identification of anomalies in the capnography waveform.
[0029] The plurality of turning points may further include a beta turning point between the expiratory plateau and the inspiratory descending limb. Determining additional turning points, such as the beta turning point, enables extraction of a greater number of features from the respiratory waveform and thereby facilitates more accurate predictions by the trained machine learning model and / or provides a better trained machine learning model. Optionally, the waveform may be segmented into an expiratory phase and an inspiratory phase based on the determined beta turning point.
[0030] Preferably, determining the plurality of turning points comprises determining a derivative of the respiratory waveform. The derivative of the respiratory waveform may be helpfully implemented in several steps of the method for various reasons. Several of these embodiments are discussed in detail below. For example, the derivative of the respiratory waveform may be used as a reference to the respiratory waveform to facilitate the determination of the turning points. Preferably, the derivative of the respiratory waveform is a first-order differential of the respiratory waveform, as first-order differentials typically include less noise than higher-order derivatives. However, higher-order derivatives (e.g., second-order differentials) may be more suitable for determining the plurality of turning points based on the analyzed respiratory waveform. Determining the plurality of turning points may also comprise determining a plurality of derivatives of the respiratory waveform. For example, the first-order differential of the respiratory waveform may be used to determine a given turning point, and the second-order differential of the respiratory waveform may be used to determine a different turning point.
[0031] Derivatives of the respiratory waveform, such as first-order differentials, can also be used to identify and, optionally, exclude abnormal respiratory waveforms, thereby preventing unnecessary processing. In some embodiments of the first or second aspects, the method further includes comparing the derivative to a template; and, when the derivative does not match the template, excluding the respiratory waveform. The template can be a different template or another type of template. Using the derivative to identify (and exclude) abnormal respiratory waveforms can also include determining whether the value of the derivative is within an expected range.
[0032] Abnormal respiratory waveforms can also be identified and optionally excluded by other means. For example, by comparing the value of the respiratory waveform itself or the value of its derivative (e.g., at random or preset points) to expected values or ranges. The template and these expected values are determined through empirical observation of typical physiology.
[0033] Preferably, determining a derivative of the respiratory waveform, such as a first-order differential of the respiratory waveform, includes applying a time-based smoothing filter, such as a Savitsky-Golay filter, to the respiratory waveform. For example, the first-order differential of the respiratory waveform is determined in conjunction with the application of a Savitsky-Golay filter. Time-based smoothing filters apply smoothing and are therefore particularly advantageous when used with noisy data. Savitsky-Golay filters are also very versatile, including the use of many window sizes and orders of polynomial fits, and have therefore been found to be effective for a wide range of different respiratory waveforms. Other examples of smoothing filters include frequency filtering (e.g., using wavelets or short-time Fourier filters) and moving average smoothing functions.
[0034] Artifacts in the respiratory waveform, such as humps, can have a significant impact on the shape of the waveform and its corresponding derivatives. In some cases, depending on the size and location of the artifact, this can lead to inconsistent feature extraction and less accurate predictions from a trained machine learning model, or less accurate training of the model. Addressing any humps is particularly important when the automated prediction method (or training method) aims to make accurate predictions on a reasonable timescale. Hums can vary significantly depending on the airway generating the capnography waveform; however, they are typically characterized by a sharp increase in pCO2 in the respiratory waveform just before the expiratory ascending limb. The increased pCO2 of the hump is less than the increased pCO2 of the expiratory plateau and may remain at the same level until the expiratory ascending limb, may partially decrease during the expiratory ascending limb, or may decrease and then fully return to the expiratory baseline pCO2 level.
[0035] Thus, preferably, determining the plurality of turning points comprises identifying a hump artifact in the respiratory waveform, and when the hump artifact is present, processing the hump artifact during determining the one or more turning points.
[0036] Hump artifacts can be handled in various ways, for example, by ignoring or removing data associated with the hump artifact, effectively subtracting the hump artifact, or adjusting the weights applied to data associated with the hump artifact when determining one or more turning points.
[0037] Preferably, identifying a hump artifact includes: performing peak detection to identify local minima of a respiratory waveform; identifying significant minima from the local minima; identifying a maximum of the respiratory waveform and / or determining a beta turning point; dividing the respiratory waveform into a first portion and a second portion, the first portion not including the maximum of the respiratory waveform, the second portion including the maximum and / or beta turning point of the respiratory waveform; searching for a hump artifact in the first portion when at least one significant minimum is identified; and / or searching for a hump artifact in the first portion of the respiratory waveform using a derivative of the respiratory waveform when no significant minimum is identified. When the derivative of the respiratory waveform is a first-order differential of the respiratory waveform, searching / identifying a hump artifact using the derivative of the respiratory waveform may include analyzing the first-order differential to identify an inflection point region of the first-order differential, and comparing the inflection point region to a preset threshold.
[0038] The maximum value of the respiratory waveform refers to the maximum amplitude of the respiratory waveform, for example, the maximum pCO2 value measured at a time point in the respiratory waveform.
[0039] The respiratory waveform is divided along the time dimension. That is, the values up to a certain point in time (the point that demarcates the first and second parts) are in the first part, and the values after that point in time are in the second part. The boundary between the first and second parts can be the maximum value of the respiratory waveform.
[0040] Alternatively, the respiratory waveform may be divided such that the first portion and the second portion are defined relative to the β turning point. That is, the respiratory waveform is divided into a first portion and a second portion, the first portion does not include the β turning point, and the second portion includes the β turning point.
[0041] It has been found that minima corresponding to hump artifacts typically have CO2 values below a certain value. Therefore, to further reduce computational cost, a hump artifact threshold can be set during hump artifact identification. For example, the hump artifact threshold can define a maximum CO2 value, where any minima with CO2 values above the hump threshold are discarded or not considered during hump artifact identification. Preferably, the hump artifact threshold is 0.5 kPa to 4 kPa. It has been found that using a threshold within this range achieves an accurate level of hump artifact for most respiratory waveforms (e.g., by ignoring minima in the expiratory plateau). Most preferably, the hump artifact threshold is 2 kPa.
[0042] Alternatively, the first portion and the second portion may be defined as non-overlapping portions, wherein the second portion includes a maximum value and / or a beta turning point of the respiratory waveform.
[0043] The beta turning point can be determined by a variety of means. Reliable and effective methods for determining the beta turning point, particularly during automation, have been found to include: performing peak detection to find local maxima of the respiratory waveform; identifying significant maxima from among the local maxima; when only a single significant maximum is identified, identifying it as the beta turning point between the expiratory plateau and the inspiratory descending limb; and, when multiple significant maxima are identified, identifying the most significant maximum and defining it as the beta turning point.
[0044] Extensive data analysis has revealed that pCO2 values at the β-turning point typically fall within a range of values, or fall below a range of values. Therefore, a preset threshold (i.e., a maximum threshold) can be defined to conserve further processing resources. This can be achieved by ignoring maxima with pCO2 values below the maximum threshold (e.g., by removing these values or setting them equal to 0 during peak detection or identification of significant maxima). This reduces the processing resources used by the method and reduces noise from values before the maximum threshold. For example, the maximum threshold can be set to values from 0.04 kPa to 5 kPa. These lower and higher threshold values account for background pressure while reliably and accurately determining the β-turning point, allowing accurate and useful features to be extracted based on the determined β-turning point. Preferably, the maximum threshold can be set between 1.5 kPa and 2.5 kPa. Most preferably, the maximum threshold is 2 kPa. The maximum threshold can be preset, or in some embodiments, it can be determined and / or adjusted by the machine learning model as it continues to determine β-turning points for increasing amounts of the respiratory waveform. This means that the machine learning model can further optimize the method's resource usage.
[0045] When no significant maxima are identified, the respiratory waveform may be excluded, as this is often due to an abnormal respiratory waveform where the extracted features are unlikely to be accurate or may not be reliably interpreted by the machine learning model (i.e., during prediction or training). Excluding the respiratory waveform prevents these outcomes at an early stage without requiring any further unnecessary processing.
[0046] Preferably, determining the turning point includes determining a delta turning point, and determining the delta turning point includes: determining a first time point at which a first-order differential of a respiratory waveform is above a delta threshold; and defining the first point as the delta turning point.
[0047] Similar to the maximum threshold, the delta threshold can be a preset threshold, or can be determined and / or adjusted by a (trained or untrained) machine learning model. Optionally, the delta threshold is configured relative to the maximum amplitude of the first order differential of the respiratory waveform. For example, the delta threshold can be between 5% and 80% of the maximum amplitude of the first order differential. Preferably, the delta threshold is between 5% and 15% of the maximum amplitude point of the first order differential. More preferably, the delta threshold is 10% of the maximum amplitude point of the first order differential. It has been found that using these ranges of delta thresholds and a specified value of 10% accurately determines the delta turning point, thereby allowing accurate and useful features to be extracted based on the determined delta turning point.
[0048] Preferably, determining the turning point includes determining a gamma turning point, and determining the gamma turning point includes: identifying a minimum in the first order differential of the respiratory waveform; and defining the gamma turning point as the first time point after the minimum at which the first order differential of the respiratory waveform is above a gamma threshold.
[0049] The first time point refers to a substantially first time point, and may be a data point near the first time point, or, for example, closely corresponding to the first time point.
[0050] Similar to the thresholds discussed previously, the gamma threshold can be a preset threshold, or can be determined and / or adjusted by a (trained or untrained) machine learning model. Optionally, the gamma threshold is configured relative to the maximum amplitude of the first order differential of the respiratory waveform. For example, the gamma threshold can be from 0 to negative 50% of the maximum amplitude point of the first order differential. Preferably, the gamma threshold can be from negative 2% to negative 10% of the maximum amplitude point of the first order differential. More preferably, the delta threshold is negative 5% of the maximum amplitude point of the first order differential. It has been found that using these ranges of gamma thresholds and a specified value of negative 5% can accurately determine the gamma turning point, so that accurate and useful features can be extracted based on the determined gamma turning point.
[0051] Preferably, determining the turning point includes: determining the α turning point, wherein determining the α turning point includes: identifying the maximum value of the first-order differential of the respiratory waveform; identifying the maximum value of the respiratory waveform and / or determining the β turning point; and defining the α turning point as the first time point at which the first-order differential of the respiratory waveform is less than the α threshold, the first time point being between the maximum value of the first-order differential and the maximum value of the respiratory waveform and / or the β turning point.
[0052] It will be apparent that the alpha turning point can be determined using the maximum value of the respiratory waveform or using the beta turning point. These two methods can be performed separately (i.e., only one can be performed to determine the alpha turning point while using fewer processing resources), or both can be performed to verify the results of the other. In the event that the determination of the beta turning point has not yet been completed, it only needs to be determined at this stage, otherwise a previously determined beta turning point can be used, or the alpha turning point can be determined without using the beta turning point. It has been found that using the beta turning point to determine the alpha turning point (and / or identify the hump artifact) is a more reliable method than using the maximum value of the respiratory waveform, and is applicable to a wider range of respiratory waveform shapes. However, while the technique using the maximum value of the respiratory waveform is reliable and accurate, it is associated with a lower computational cost.
[0053] Similar to the thresholds discussed previously, the alpha threshold can be a preset threshold, or can be determined and / or adjusted by a (trained or untrained) machine learning model. Optionally, the alpha threshold is configured relative to the maximum amplitude of the first order differential of the respiratory waveform. For example, the alpha threshold can be from 0 to 80% of the maximum amplitude point of the first order differential. Preferably, the alpha threshold can be from 10% to 20% of the maximum amplitude point of the first order differential. More preferably, the alpha threshold is 15% of the maximum amplitude point of the first order differential. It has been found that using these ranges of alpha thresholds and a specified value of 15% can accurately determine the alpha turning point, so that accurate and useful features can be extracted based on the determined alpha turning point.
[0054] In some embodiments of the method, the alpha threshold can be adjusted to reduce the impact of noise on the determination of the alpha turning point and improve the processing resource efficiency of the method. Therefore, determining the alpha turning point can further include: when there is no point between the maximum value of the first-order differential of the respiratory waveform and the maximum value of the respiratory waveform that is less than the alpha threshold, increasing the alpha threshold. The alpha threshold can be increased by a predetermined amount. For example, the alpha threshold can be increased to 5% of the maximum amplitude point of the first-order differential of the respiratory waveform.
[0055] After increasing the alpha threshold, the value of the first-order differential between the maximum value of the first-order differential and the maximum value of the respiratory waveform and / or the β turning point (i.e., between the time points corresponding thereto) can be compared with the increased alpha threshold to determine and define the alpha turning point. If no point meets this definition, the alpha threshold can be increased again, and this process can be iteratively performed until the alpha turning point is determined. The amount by which the alpha threshold is increased can vary between iterations.
[0056] To reduce unnecessary processing resources used during the method, upper limits may be placed on the number of iterations performed (and therefore the number of times the alpha threshold is raised) and / or the value of the alpha threshold itself, and respiratory waveforms excluded when one of these limits is reached.
[0057] The alpha turning point may also be determined using other alternative methods. One such alternative method for determining the alpha turning point includes calculating a line between the delta turning point and a maximum value of the respiratory waveform, or calculating a line between the delta turning point and the beta turning point; and defining the alpha turning point based on the distance between the respiratory waveform and the calculated line.
[0058] For example, the α turning point can be defined as a point on the respiratory waveform between the δ turning point and the maximum value of the respiratory waveform (or the β turning point), i.e., a point farthest from a line calculated between the δ turning point and the maximum value of the respiratory waveform (or between the δ turning point and the β turning point), where the distance between the respiratory waveform and the calculated line is measured using another line perpendicular to the (first) calculated line (between the δ turning point and the maximum value of the respiratory waveform or the β turning point).
[0059] This can be implemented as an alternative method to using an alpha threshold to determine the alpha turning point, or used as a complementary method to determine the alpha turning point to validate another method.
[0060] In an alternative embodiment of the algorithmic scheme above, determining the turning point may include: applying a trained machine learning model to a set of discrete samples of a respiratory waveform, the respiratory waveform representing a complete breath, wherein the machine learning model is configured to classify each sample into one of a plurality of output categories, each category representing a region of the respiratory waveform, and wherein the machine learning model is trained by: obtaining a label associated with each discrete sample in a plurality of respiratory waveforms, each respiratory waveform being represented by a set of samples representing a complete breath, and each label indicating that the sample corresponds to one of a plurality of output categories; and, training the machine learning model based on the labels and the samples to learn to classify the set of samples representing a complete breath into one of the plurality of output categories. This approach may provide consistency for feature engineering, but may be difficult to account for all types of breaths, and accuracy will be proportional to the amount of input data for the labels.
[0061] Optionally, the method further includes extracting features from multiple respiratory waveforms recorded from the same user; and determining parameters of the extracted features; wherein the variability of the extracted features is further used to apply a trained machine learning model to classify capnography (and, in some embodiments, to predict cardiopulmonary disease), or wherein the variability of the extracted features is further used to train a machine learning model to create a classification function (in some embodiments, the classification function is a prediction function). The parameter may, for example, be the variability of the extracted feature over a time period of the recorded waveforms, or further, the parameter may be a metric associated with the time-varying feature, such as a mean, median, standard deviation, distribution, or other suitable similarity score, such as cosine similarity or distance. The variability of the extracted features may be determined using two or more respiratory waveforms recorded from the same user. More preferably, the multiple respiratory waveforms recorded from the same user include at least three respiratory waveforms recorded from the same user. It has been found that determining the variability from at least three respiratory waveforms provides accurate classification and / or an accurate classification function.
[0062] The time period used to determine the variability of the extracted features can vary depending on the features being examined and / or the cardiopulmonary disease that the machine learning model is predicting (or being trained to predict). The time period over which the respiratory waveform is recorded can be, for example, multiple hours, multiple days, multiple weeks, or multiple months. For some cardiopulmonary diseases, the variation of the extracted features over time is a particularly important indicator when predicting whether the carbon dioxide waveform is produced by the cardiopulmonary airway suffering from the disease. For example, the time period can include two or more days, preferably the time period includes five or more days, and more preferably the time period includes ten or more days. For some cardiopulmonary diseases, such as obstructive airway diseases (such as asthma) that are characterized by airway changes, the time period over which the respiratory waveform is recorded is less important, and accurate classification (or accurate classification function) can be achieved by determining the variation without considering this time period.
[0063] Preferably, multiple respiratory waveforms are recorded at substantially regular intervals within the time period. For example, when the time period is five days, the multiple respiratory waveforms include a first respiratory waveform obtained from a first capnogram recorded on the first day of the five days, a second respiratory waveform obtained from a second capnogram recorded on the second day of the five days, a third respiratory waveform obtained from a third capnogram recorded on the third day of the five days, and so on. More preferably, capnograms are recorded two or more times per day to assess respiratory waveform variability and features extracted throughout the day. For example, morning and evening variability has been found to be particularly effective for classification related to asthma.
[0064] In certain instances, when predicting whether a capnogram is produced by a user with cardiopulmonary disease, multiple capnograms should each be obtained from a single user's recorded capnograms.
[0065] Various features of the respiratory waveform can be extracted using the determined turning points. For example, the extracted features can include at least one of the following: time and / or CO2 pressure at the turning point or any other point on the respiratory waveform, end-tidal CO2, the angle between linear fits on either side of the turning point, the duration of the phase between the turning points, the ratio of the angle between the linear fits to the duration of the phase, the minimum, maximum, mean, median, and / or total CO2 pressure during the phase between the turning points, respiratory rate, and coefficients of a line fitted to the respiratory waveform using the turning points.
[0066] Further examples include: a hyperbolic tangent fit, a delta angle, an alpha angle, a beta angle, a gamma angle, a delta angle starting point, a delta angle ending point, an alpha angle starting point, an alpha angle ending point, a beta angle starting point, a beta angle ending point, a gamma angle starting point, a gamma angle ending point and the times and CO2 values at these points, the minimum CO2 in the inspiratory baseline of the respiratory waveform, the minimum CO2 in the expiratory ascending limb of the respiratory waveform, the ratio of the alpha angle to the duration of the expiratory phase, the CO2 at the center of the gamma angle, the coefficients of a quadratic fit of the expiratory plateau, the time difference between any identified turning points, respiratory patterns, and measures of disorder (e.g., entropy).
[0067] Features related to breathing patterns are particularly useful for classifying capnograms for breathing pattern disorders. These can be detected (e.g., using autocorrelation or frequency domain identifiers) by assessing the periodicity of the breathing waveform and comparing the variability of breathing.
[0068] The delta angle start point refers to the capnogram value at the delta angle start point, such as the time and / or pCO2 value at the delta angle start point. The same applies to the start and end points of the alpha, beta, and gamma angles.
[0069] Extracting features of a respiratory waveform using a turning point may include determining an angle of the turning point. Determining the angle of the turning point may include fitting a first linear function and a second linear function to adjacent phases on either side of the turning point, and measuring the angle between the first linear function and the second linear function. For example, if the angle of the α turning point is to be determined, fitting the first and second linear functions to adjacent phases on either side of the α turning point means fitting the first linear function to the expiratory ascending limb and fitting the second linear function to the expiratory plateau.
[0070] Extracting features of the respiratory waveform using turning points can also include fitting a quadratic function to the expiratory plateau and determining one or more coefficients of the quadratic function. This provides an indication of the "flatness" of the expiratory plateau, which can be used to help determine features of the airway associated with the respiratory waveform. Preferably, the quadratic function is fitted across the entire length of the expiratory plateau, from the α turning point to the β turning point.
[0071] As a further example, extracting features of the respiratory waveform using turning points can include fitting a hyperbolic tangent function to the expiratory ascending limb and / or the inspiratory descending limb and determining one or more coefficients of the hyperbolic tangent function. These coefficients include, for example, the width and horizontal shift parameters of the fitted hyperbolic tangent function. Preferably, the hyperbolic tangent function is fitted across the entire length of the expiratory ascending limb (i.e., between the δ and α turning points) and the entire length of the inspiratory descending limb (i.e., between the β and γ turning points).
[0072] When used in the methods of the present invention, some machine learning or statistical inference models have been found to be particularly effective (e.g., providing high-accuracy predictions). Therefore, preferably, the machine learning model includes at least one of the following: logistic regression, gradient boosted decision trees, support vector machines, ensemble methods (e.g., AdaBoost), and random forests. This is not an exhaustive list of models.
[0073] In some embodiments of the first aspect, the trained machine learning model is further configured to output an indication of the importance of the extracted features that led to the classification of the capnogram. That is, in some embodiments, predicting cardiopulmonary disease. The method may include the step of outputting the indication. In this manner, the indication can be used to evaluate the trained machine learning model and to explain the predictions by providing additional context for the predictions made.
[0074] According to a third aspect, a method for classifying one or more carbon dioxide waveforms generated by a user is provided, the method comprising: obtaining multiple respiratory waveforms from the one or more carbon dioxide waveforms generated by the user; normalizing the durations of the multiple respiratory waveforms to generate multiple standardized respiratory waveforms; generating an average respiratory waveform from the multiple standardized respiratory waveforms (preferably using a generalized additive model, GAM); extracting features from the average respiratory waveform; and applying a trained machine learning model to the extracted features, wherein the trained machine learning model is configured to output a classification from the one or more carbon dioxide waveforms based on the features.
[0075] Similar to the first aspect, classification of one or more capnograms can be used to predict cardiopulmonary disease. These embodiments can be considered as methods for predicting whether one or more capnograms are generated by a user with cardiopulmonary disease, wherein a trained machine learning model is configured to predict cardiopulmonary disease based on features.
[0076] That is, an embodiment of the third aspect includes a method for predicting whether one or more carbon dioxide waveform graphs are generated by a user suffering from cardiopulmonary disease, the method comprising: obtaining multiple respiratory waveforms from one or more carbon dioxide waveform graphs generated by the user; standardizing the durations of the multiple respiratory waveforms to generate multiple standardized respiratory waveforms; generating an average respiratory waveform from the multiple standardized respiratory waveforms (preferably using a generalized additive model, GAM); extracting features from the average respiratory waveform; and applying a trained machine learning model to the extracted features, wherein the trained machine learning model is configured to predict cardiopulmonary disease based on the features.
[0077] In this way, a single smoothed average respiratory waveform is generated, from which features can then be extracted. Extracting features from the average respiratory waveform can be performed using the same techniques discussed above in the embodiments of the first and second aspects, instead of, or in addition to, extracting features from individual respiratory waveforms. Similarly, determination of turning points and identification of artifacts can also be performed using the same methods discussed above in the first and second aspects. Those skilled in the art will recognize that the method steps related to determining turning points and artifacts are applicable to the third aspect.
[0078] Optionally, the duration and amplitude of the plurality of respiratory waveforms are normalized to generate a plurality of normalized respiratory waveforms. That is, the normalized respiratory waveforms have been normalized with respect to duration and amplitude.
[0079] The third aspect of the present invention can be combined with any one of the first aspect or the second aspect to improve these aspects, and can also be separated from each of these aspects so that it can be implemented separately from these aspects.
[0080] According to a fourth aspect, a method for training a machine learning model to learn an indicator of a capnography waveform is provided, the method comprising: for each average respiratory waveform, generating a plurality of average respiratory waveforms by the following steps: obtaining a plurality of respiratory waveforms from one or more capnograms generated by a user; normalizing the durations of the plurality of respiratory waveforms to generate a plurality of normalized respiratory waveforms; and generating an average respiratory waveform from the plurality of normalized respiratory waveforms (preferably using a generalized additive model, GAM); Extracting features from each of a plurality of average respiratory waveforms; obtaining a label for each average respiratory waveform, the label indicating a capnogram classification of the corresponding one or more capnograms; and training a machine learning model to learn capnogram indicators using the extracted features of the plurality of average respiratory waveforms and the corresponding labels to create a classification function.
[0081] Similar to the second aspect, in some embodiments of the fourth aspect, a machine learning model is trained to learn indicators of cardiopulmonary disease (i.e., indicators of carbon dioxide waveforms generated by users with cardiopulmonary disease), wherein the carbon dioxide waveform classification of the corresponding one or more carbon dioxide waveforms indicates whether the one or more carbon dioxide waveforms are from a diseased respiratory tract and / or circulatory system.
[0082] That is, an embodiment of the fourth aspect includes a method for training a machine learning model to learn an indicator of a cardiopulmonary disease, the method comprising: for each average respiratory waveform, generating a plurality of average respiratory waveforms by the following steps: obtaining a plurality of respiratory waveforms from one or more capnograms generated by a user; normalizing the durations of the plurality of respiratory waveforms to generate a plurality of normalized respiratory waveforms; and Generate a mean respiratory waveform from multiple normalized respiratory waveforms (preferably using a generalized additive model, GAM); Extracting features from each of a plurality of average respiratory waveforms; obtaining a label for each average respiratory waveform, the label indicating whether the corresponding one or more carbon dioxide waveforms are from a diseased respiratory tract and / or circulatory system; and training a machine learning model to learn indicators of cardiopulmonary disease using the extracted features of the plurality of average respiratory waveforms and the corresponding labels to create a prediction function.
[0083] Optionally, the duration and amplitude of the plurality of respiratory waveforms are normalized to generate a plurality of normalized respiratory waveforms. That is, the normalized respiratory waveforms have been normalized with respect to duration and amplitude.
[0084] The fourth aspect of the present invention can be combined with any one of the first to third aspects to improve these aspects, and can also be separated from each of these aspects so that it can be implemented separately from these aspects.
[0085] Although several of the techniques relating to the third and fourth aspects of generating averaged waveforms are described in the context of predicting whether one or more carbon dioxide waveforms are produced by a user with a cardiopulmonary disease, those skilled in the art will recognize that these averaged waveform techniques are not limited to use with cardiopulmonary diseases or predictive methods and may be implemented with any similarly shaped waveform.
[0086] Normalizing the amplitudes of the plurality of respiratory waveforms may include extracting an end-tidal CO2 value from each respiratory waveform; and adjusting the amplitude of each of the plurality of respiratory waveforms so that each of the plurality of respiratory waveforms has the same end-tidal CO2 value. The amplitudes of the plurality of respiratory waveforms may be normalized based on various points of the individual respiratory waveforms, such as a maximum or other inflection point, such as an alpha inflection point, of the respiratory waveforms. However, it has been found that normalizing the waveforms based on the end-tidal CO2 value consistently provides a normalized waveform well-suited for use with the GAM to produce a smooth and accurate average respiratory waveform.
[0087] Extracting the end-tidal CO2 value from each respiratory waveform may include, for each of the plurality of respiratory waveforms, determining: a beta turning point between an expiratory plateau and an inspiratory descending limb.
[0088] Preferably, obtaining the plurality of respiratory waveforms comprises: dividing one or more incoming capnograms into a plurality of capnogram portions, each portion representing a single respiratory waveform corresponding to a single respiratory cycle; and for each of the plurality of capnogram portions: determining a delta turning point between an expiratory baseline and an expiratory ascending limb; determining a gamma turning point between an inspiratory descending limb and the inspiratory baseline; and extracting the portion of the capnogram between the delta turning point and the gamma turning point to generate the respiratory waveform.
[0089] Generating an averaged respiration waveform from a plurality of normalized respiration waveforms may include: identifying whether any of the plurality of normalized respiration waveforms is abnormal; and excluding the abnormal normalized respiration waveform when generating the averaged respiration waveform (preferably by using a GAM). Excluding abnormal respiration waveforms increases the usefulness of the resulting averaged respiration waveform generated using the normalized respiration waveforms. Abnormal respiration waveforms may also be identified without normalizing the respiration waveforms; however, normalizing the respiration waveforms may improve the accuracy of any subsequent abnormality detection.
[0090] Preferably, identifying whether any one of the plurality of standardized respiratory waveforms is abnormal comprises: interpolating the standardized respiratory waveforms so that all the standardized respiratory waveforms have the same number of data points; comparing the data points of each interpolated respiratory waveform with the data points of the other interpolated respiratory waveforms; and identifying any interpolated respiratory waveform having one or more abnormal data points as abnormal.
[0091] The interpolated respiration waveform refers to a waveform generated by interpolating a normalized respiration waveform. This method has been found to be less computationally demanding than other methods of excluding abnormal respirations, such as standard anomaly detection algorithms.
[0092] Preferably, the method of the first aspect and the third aspect further comprises obtaining a capnogram from the user, wherein the capnogram comprises a respiratory waveform.
[0093] According to a fifth aspect, there is provided a method of diagnosing a cardiopulmonary disease, the method comprising the method according to any embodiment of the first aspect or the third aspect, and further comprising: using classification of a capnography waveform to diagnose the presence of the cardiopulmonary disease.
[0094] Obviously, the classification of the capnogram is used to diagnose the presence of cardiopulmonary disease, and this diagnosis is made remotely from the user who generated the capnogram corresponding to the classification.
[0095] According to a sixth aspect, a device is provided, which is configured to perform the method described in any one of the first to fifth aspects.
[0096] According to a seventh aspect, a device for classifying a carbon dioxide waveform is provided, the device comprising: an acquisition module configured to obtain a respiratory waveform from the carbon dioxide waveform; a determination module configured to determine one or more turning points of the respiratory waveform, wherein the turning points include: a δ turning point between the expiratory baseline and the expiratory ascending limb, a γ turning point between the inspiratory descending limb and the inspiratory baseline, and an α turning point between the expiratory ascending limb and the expiratory platform; the device further comprises an extraction module configured to extract features of the respiratory waveform using the turning points; and a prediction module configured to apply a trained machine learning model to the extracted features, wherein the trained machine learning model is configured to classify the carbon dioxide waveform according to the features.
[0097] An embodiment of the seventh aspect includes a device for predicting whether a carbon dioxide waveform graph is generated by a user suffering from cardiopulmonary disease, the device comprising: an acquisition module configured to acquire a respiratory waveform from the carbon dioxide waveform graph; a determination module configured to determine one or more turning points of the respiratory waveform, wherein the turning points include: a δ turning point between the expiratory baseline and the expiratory ascending limb, a γ turning point between the inspiratory descending limb and the inspiratory baseline, and an α turning point between the expiratory ascending limb and the expiratory platform; the device further comprises an extraction module configured to extract features of the respiratory waveform using the turning points; and a prediction module configured to apply a trained machine learning model to the extracted features, wherein the trained machine learning model is configured to predict cardiopulmonary disease based on the features.
[0098] According to an eighth aspect, a device for training a machine learning model to learn indicators of a carbon dioxide waveform graph is provided, the device comprising: an acquisition module configured to acquire a respiratory waveform from each of a plurality of carbon dioxide waveform graphs, and further configured to acquire a label for each carbon dioxide waveform graph, the label indicating a classification of the carbon dioxide waveform graph; a determination module configured to determine one or more turning points of the respiratory waveform, wherein the turning points comprise: a delta turning point between an expiratory baseline and an expiratory ascending limb, a gamma turning point between an inspiratory descending limb and an inspiratory baseline, and an alpha turning point between an expiratory ascending limb and an expiratory platform; an extraction module configured to extract features of the respiratory waveform using the one or more turning points; and a training module configured to train the machine learning module using the extracted features of the plurality of respiratory waveforms and the labels corresponding to the respective carbon dioxide waveform graphs to learn indicators of the carbon dioxide waveform graph to create a classification function.
[0099] An embodiment of the eighth aspect includes a device for training a machine learning model to learn indicators of a carbon dioxide waveform graph generated by a user with cardiopulmonary disease, the device comprising: an acquisition module configured to acquire a respiratory waveform from each of a plurality of carbon dioxide waveform graphs, and further configured to acquire a label for each carbon dioxide waveform graph, the label indicating whether the carbon dioxide waveform graph is from a diseased respiratory tract and / or circulatory system; a determination module configured to determine one or more turning points of the respiratory waveform, wherein the turning points include: a delta turning point between an expiratory baseline and an expiratory ascending limb, a gamma turning point between an inspiratory descending limb and an inspiratory baseline, and an alpha turning point between an expiratory ascending limb and an expiratory platform; an extraction module configured to extract features of the respiratory waveform using one or more turning points; and a training module configured to train the machine learning module using the extracted features of the plurality of respiratory waveforms and the corresponding labels of the respective carbon dioxide waveform graphs to learn indicators of cardiopulmonary disease to create a prediction function.
[0100] According to a ninth aspect, a device for classifying one or more carbon dioxide waveform graphs generated by a user is provided, the device comprising: an acquisition module configured to acquire a plurality of respiratory waveforms from the one or more carbon dioxide waveform graphs generated by the user; a normalization module configured to normalize the durations of the plurality of respiratory waveforms to generate a plurality of standardized respiratory waveforms; a generation module configured to generate an average respiratory waveform from the plurality of standardized respiratory waveforms (preferably using a generalized additive model, GAM); an extraction module configured to extract features from the average respiratory waveform; and a prediction module configured to apply a trained machine learning model to the extracted features, wherein the trained machine learning model is configured to output a classification of the one or more carbon dioxide waveform graphs based on the features.
[0101] An embodiment of the ninth aspect includes a device for predicting whether one or more carbon dioxide waveform graphs are generated by a user suffering from cardiopulmonary disease, the device comprising: an acquisition module configured to acquire multiple respiratory waveforms from one or more carbon dioxide waveform graphs generated by the user; a normalization module configured to normalize the duration of the multiple respiratory waveforms to generate multiple standardized respiratory waveforms; a generation module configured to generate an average respiratory waveform from the multiple standardized respiratory waveforms (preferably using a generalized additive model, GAM); an extraction module configured to extract features from the average respiratory waveform; and a prediction module configured to apply a trained machine learning model to the extracted features, wherein the trained machine learning model is configured to predict cardiopulmonary disease based on the features.
[0102] According to a tenth aspect, a device for training a machine learning model to learn an indicator of a carbon dioxide waveform graph is provided, the device comprising: an acquisition module configured to acquire a plurality of respiratory waveforms from one or more carbon dioxide waveform graphs generated by a user; a normalization module configured to normalize the durations of the plurality of respiratory waveforms to generate a plurality of standardized respiratory waveforms; a generation module configured to generate an average respiratory waveform from the plurality of standardized respiratory waveforms (preferably using a generalized additive model, GAM); an extraction module configured to extract features from each of the plurality of average respiratory waveforms; wherein the acquisition module is further configured to acquire a label for each average respiratory waveform, the label indicating a carbon dioxide waveform classification of the corresponding one or more carbon dioxide waveform graphs; and wherein the device further comprises a training module configured to train the machine learning model to learn an indicator of the carbon dioxide waveform graph using the extracted features of the plurality of average respiratory waveforms and the corresponding labels to create a classification function.
[0103] An embodiment of the tenth aspect includes an apparatus for training a machine learning model to learn indicators of cardiopulmonary disease, the apparatus comprising: an acquisition module configured to acquire a plurality of respiratory waveforms from one or more carbon dioxide waveform graphs generated by a user; a normalization module configured to normalize the durations of the plurality of respiratory waveforms to generate a plurality of standardized respiratory waveforms; a generation module configured to generate an average respiratory waveform from the plurality of standardized respiratory waveforms (preferably using a generalized additive model, GAM); an extraction module configured to extract features from each of the plurality of average respiratory waveforms; wherein the acquisition module is further configured to obtain a label for each average respiratory waveform, the label indicating whether the corresponding one or more carbon dioxide waveform graphs are from a diseased respiratory tract and / or circulatory system; and wherein the apparatus further comprises a training module configured to use the extracted features and corresponding labels of the plurality of average respiratory waveforms to train the machine learning model to learn indicators of cardiopulmonary disease to create a prediction function.
[0104] In addition to the above examples, the capnogram classification can include a probability value corresponding to the severity of a respiratory disease. In other words, the output classification can include a probability value. Similarly, the capnogram classification can include a value indicating the likelihood of the capnogram having a respiratory disease severity level. The probability value can indicate the confidence that the capnogram has a preset severity value, i.e., a mild severity level. The probability value can indicate whether the capnogram is likely to be mild or severe, and thus can represent where the capnogram falls within a range between mild and severe. In other words, it is a continuous index quantifying the severity of a disease (e.g., COPD), where the model is configured to output a continuous prediction probability.
[0105] Additionally, the classification function may be configured to predict the severity of respiratory disease.
[0106] Additionally, a machine learning model can be trained based on capnograms labeled with corresponding severity levels of respiratory disease. Each capnogram or respiratory waveform can be labeled with a corresponding severity level, either as a continuous variable or from a pre-set severity group. Preferably, each capnogram can be labeled as mild or severe, so that the probability value output by the model during implementation indicates a range of disease severity.
[0107] According to further embodiments, a machine learning model can be trained and configured to predict the likelihood that a capnogram is associated with a cardiopulmonary disease. Again, we note that "disease" also refers to a subtype of the disease. The cardiopulmonary disease can be selected from the group consisting of COPD, asthma, asthma-COPD overlap syndrome (ACOS), small airway disease, the chronic bronchitis subtype of COPD, and the emphysema subtype of COPD. Therefore, it should be understood that a disease can refer to both a disease and a subtype of the disease, but for the sake of clarity and brevity, we will not repeat the phrase "disease or subtype."
[0108] Classification can be binary or multi-class. A machine learning model can be trained using labels that belong to (or indicate) a category indicating that a capnogram is associated with a cardiopulmonary disease. Labels can further belong to multiple categories, each category indicating a capnogram associated with a cardiopulmonary disease, each category corresponding to a corresponding disease. Labels can further belong to a category indicating that a capnogram is not associated with a cardiopulmonary disease. In this way, a binary classifier can be configured to predict the likelihood of a disease or the absence of a disease, or the likelihood of one of multiple diseases. This problem can also be extended to multi-class classification to predict one of multiple diseases.
[0109] The machine learning model can be configured to output a likelihood that the capnogram waveform is associated with cardiopulmonary disease. The method can also include stratifying the output by comparing the likelihood to a set of thresholds, with each stratum representing the risk of the capnogram waveform being associated with cardiopulmonary disease. This improves the clinical utility of the model, particularly when combined with an interpretable model.
[0110] According to an eleventh aspect, a computer-readable medium comprising instructions is provided. When the instructions are executed by a processor, the instructions cause the processor to: execute the method of the first, second, third, fourth and / or fifth aspects above.
[0111] BRIEF DESCRIPTION OF THE DRAWINGS
[0112] Embodiments of the invention are described below, by way of example only, with reference to the accompanying drawings, in which: Figure 1A Respiratory waveforms from a healthy individual are shown; Figure 1B Respiratory waveforms from patients with cardiopulmonary disease are shown; Figure 1C The respiratory waveform including artifacts is shown; Figure 2 An exemplary capnography waveform is shown; Figure 3 An idealized respiratory waveform is shown; Figure 4A method for predicting cardiopulmonary disease according to an example of the present disclosure is shown; 5A to 5D Different respiratory waveforms and the first-order differential of the respiratory waveform are shown; Figures 6A to 6C A schematic process flow chart for identifying hump artifacts, turning points, and features of a respiratory waveform is shown; Figure 7 Methods for training a machine learning model to learn metrics from capnograms produced by users with cardiopulmonary diseases are shown; Figure 8 A schematic diagram showing an example of the present disclosure applied to three components; and Figures 9A to 9C The schematic diagram shows the details Figure 8 The three parts. Detailed Description of the Invention
[0114] Specific embodiments of the present invention will be discussed in detail below, which involve using a trained machine learning model to predict whether a carbon dioxide waveform is generated by a user with a cardiopulmonary disease, and training a machine learning model to learn indicators of a carbon dioxide waveform generated by a user with a cardiopulmonary disease.
[0115] Recent hardware innovations (such as those described in WO2017174983 and WO201915800, the contents of each of which are incorporated herein by reference) have led to the ability to reliably and non-invasively measure CO2 concentration in the mouth, closer to the airways and with high temporal resolution. This allows for greater flexibility in lung health assessment, providing more insights than previously realized.
[0116] One object of the present invention is to build a classifier that distinguishes between patients with and without cardiopulmonary disease, such as the capnogram produced by lungs with COPD or asthma, or the capnogram produced by patients susceptible to heart failure. Another object is to provide interpretability of the object so that clinicians can understand the reasons leading to the prediction and efficient computational processes suitable for deployment at the edge, such as enabling the implementation phase to use trained models at the diagnostic site without requiring cloud communication for data processing, feature extraction, and prediction. Alternatively, for example, the prediction method can be implemented on a chip, i.e., on a handheld device, or the task can be distributed between a handheld device and a local computer.
[0117] Previous academic work has contributed significantly to the understanding of how capnometry can be used to make diagnostic decisions. However, despite some impressive results from manual processing, medical administrators are still awaiting interpretable, practical clinical implementations. The method proposed here has particular utility in triaging patients with a variety of confounding conditions, such as asthma, COPD, pulmonary fibrosis, CHF, motor neuron disease (MND), breathing pattern disorders, and pneumonia.
[0118] Figures 1A to 1C Three examples of different respiratory waveforms are shown, where the respiratory wave plots the partial pressure of carbon dioxide (pCO2) as a function of time as a patient takes a single breath by exhaling and then inhaling. Unless otherwise specified, the terms partial pressure of carbon dioxide, concentration of carbon dioxide, and carbon dioxide are used interchangeably throughout this specification. Figure 1A shows an exemplary respiratory waveform resulting from the breathing of a healthy patient, Figure 1B shows an exemplary respiratory waveform resulting from the breathing of a patient with a cardiopulmonary disease (in this case COPD), and Figure 1C An exemplary respiratory waveform is shown for a patient whose breathing produces a hump-shaped artifact in the respiratory waveform, the hump artifact occurring when the patient begins to exhale.
[0119] Deoxygenated but CO2-rich blood travels through the pulmonary vasculature to the alveoli where gas exchange occurs. Perfusion exchanges O2 from the ventilation air for carbon dioxide in the blood to facilitate respiration. Diaphragmatic relaxation then allows the thoracic cavity volume to decrease during exhalation, thereby increasing pressure and forcing CO2 out through the bronchioles and into the environment. Although many embodiments use volumetric capnometry, the present disclosure focuses on temporal capnometry, the purpose of which is to measure changes in the partial pressure of carbon dioxide (pCO2) over time. It should be understood that aspects of the present invention are also applicable to volumetric capnometry. In an ideal gas mixture, the total pressure is related to the partial pressures of the individual gases by
[0120] Among them is gas The partial pressure of
[0121] in For any macroscopic volume of gas The number of molecules, and is the number of molecules in the entire gas mixture in the same volume. At sea level, the atmospheric pressure is , where background , which depends on the environment.
[0122] Typically, when a patient exhales, the CO2 concentration at the mouth increases sharply from baseline and typically reaches a plateau as diffusion begins to compensate. Inspiration then brings atmospheric CO2 back to the mouth, reducing the concentration back to baseline. The endpoint of the plateau phase is called end-tidal CO2 (ETCO2) and is an important biomarker for anesthesiologists.
[0123] like Figure 1B As shown in Figure 2, patients with obstructive cardiopulmonary disease often exhibit a more "shark fin" shape in their capnography respiratory waveform. During expiration, damaged and / or inflamed alveoli and / or airways may limit perfusion, may reduce the elasticity of the alveolar walls, and may alter the physiology of the airways, changing the force with which gas can be exhaled, which generally results in a flatter expiratory ascending limb, a larger delta angle, and a steeper expiratory plateau.
[0124] from Figures 1A to 1C It can be clearly seen that the respiratory waveform varies depending on the patient's condition and, therefore, can be examined to predict the presence of cardiopulmonary disease. The respiratory waveform varies in a variety of specific ways, and many different features can be extracted and examined from the waveform to predict disease. These waveform variations and features are discussed in more detail below.
[0125] As mentioned above, recent innovations have made it possible to obtain reliable, highly sensitive pCO2 recordings directly from the mouth using specialized handheld devices. However, it should be understood that the innovations disclosed herein are not limited to data obtained from such handheld devices, and the description is provided for context only. The patient performs normal Cheyne-Stokes breathing through the mouthpiece for approximately 75 seconds, resulting in a noninvasive and effortless experience.
[0126] Carbon dioxide strongly absorbs electromagnetic radiation at wavelengths of 4.3 μm or 15 μm, where the energy causes vibrations in its molecular bonds, which are then re-emitted in different directions and at different wavelengths. An emitter in the handheld device transmits photons of either of these wavelengths through the airway. Unabsorbed photons are detected by a sensor whose output depends on the partial pressure of carbon dioxide within the airway volume. The handheld device differs from traditional capnography in that it lacks a reference channel; instead, it self-calibrates upon startup based on the background CO2 level at the time of use. The sensor also has a fast response time, due to a combination of the velocity of respiratory airflow through the sampling volume and its proximity to the mouth, making it unaffected by differences in flow velocity when sampling away from the mouth due to wall friction (as is the case with other technologies).
[0127] The resulting output is sampled at 10kHz and reported at 50Hz, which is much higher than any other capnography monitor on the market. The anonymized data is then automatically and securely uploaded to a cloud platform via a mobile network. From there, it is run through a processing pipeline as described herein according to aspects of the present disclosure and is available for subsequent analysis. An example of a complete raw breath record is shown in Figure 2 middle.
[0128] Figure 2 An example of a tidal ventilatory capnogram is shown, which also plots the partial pressure of carbon dioxide (pCO2) as a function of time as the patient breathes. Figures 1A to 1C and Figure 2 It can be clearly seen from the comparison that Figure 2 The capnogram shows over a dozen complete breaths, as well as partial breaths at the beginning and end of the capnogram. The number of respiratory waveforms in the capnogram will vary depending on the length of time the patient's breathing is measured.
[0129] Thus, the input to the process of the present disclosure can be a set of digitized samples at a specific sampling rate, where each digital sample represents the amplitude of pCO2 at a specific point in time. The samples can be represented as a set of vectors, each with a corresponding time index. This time series data represents the input to the method.
[0130] Stage segmentation
[0131] According to the present disclosure, a classifier has been developed that utilizes several key geometric elements identified in common in respiratory waveforms to facilitate feature extraction and predict cardiopulmonary disease based on corresponding biomarkers. These geometric elements or biomarkers have not been reliably and accurately identified by computers to date, as will be explained in more detail below. Figure 3 An idealized respiratory waveform is shown, highlighting several of these geometric features.
[0132] The ideal waveform shown can be divided into five linear segments or phases, which are, from left to right as time progresses, the expiratory baseline P1, the expiratory ascending limb P2, the expiratory plateau P3, the inspiratory descending limb P4a, and the inspiratory baseline P4b. The expiratory baseline P1 is also called phase 1, the expiratory ascending limb P2 is also called phase 2, the expiratory plateau P3 is also called phase 3, the inspiratory descending limb P4a is also called phase 4a, and the inspiratory baseline P4b is also called phase 4b. Typically, when a patient exhales, the CO2 concentration at the mouth increases sharply from the expiratory baseline P1 and reaches a plateau (expiratory plateau P3) as diffusion begins to compensate for the increased CO2 concentration. Inhalation then brings atmospheric CO2 back to the patient's mouth, and the concentration decreases back to baseline (inspiratory baseline P4b). From Figure 2 As evident from the capnography waveforms, the inspiratory baseline P4b of the first respiratory waveform can extend into or serve as the expiratory baseline P1 of the second respiratory waveform, and vice versa. Therefore, the expiratory baseline P1 and the inspiratory baseline P4b are referred to herein as a single respiratory cycle, i.e., a single waveform representing one respiratory cycle, to aid interpretation of individual respiratory waveform analysis.
[0133] As will be discussed below, determining the location of at least one of the turning points between these five segments allows for reliable and efficient feature extraction of the respiratory waveform, but doing so is not straightforward. The delta turning point δ is the point on the respiratory waveform where the expiratory baseline P1 and the expiratory ascending limb P2 intersect. The alpha turning point α is the point on the respiratory waveform where the expiratory ascending limb P2 intersects the expiratory plateau P3. The beta turning point β is the point on the respiratory waveform where the expiratory plateau P3 intersects the inspiratory descending limb P4a. The gamma turning point γ is the point on the respiratory waveform where the inspiratory descending limb P4a intersects the inspiratory baseline.
[0134] Although Figure 3 The respiratory waveform in is idealized with linear segments, but this is not how the respiratory waveform appears in reality. Figure 1C Artifacts in the pre-expiratory rise (the P2 hump, as shown) are common in capnography and complicate analysis (e.g., determining the location of turning points) and feature extraction. Such artifacts are particularly problematic when automating processing or using machine learning models to facilitate prediction. Many other types of artifacts are also found in respiratory waveforms and must also be addressed to ensure that prediction methods are robust to as many types of artifacts as possible.
[0135] Noise in the measured pCO2 signal also makes it more difficult to accurately determine the turning point, as does the overall shape of the waveform. For example, it will be appreciated that Figure 1A The alpha turning point α ratio in the "square wave" waveform Figure 1BThe same alpha turning point α in the "shark fin" waveform is more clearly defined and therefore easier to determine accurately.
[0136] Each of these issues (e.g., artifacts and noise) can become more pronounced in capnograms generated at higher measurement frequencies. However, if these issues are properly addressed when determining turning points and extracting features from the respiratory waveform, the extracted features are more accurate, and therefore the resulting predictions are more accurate.
[0137] Early implementations of computer-assisted methods consisted of a manual initial segmentation step that identified Figure 3 For each phase depicted in the image, the features include the maximum and minimum CO2 values, as well as the respiratory phase gradient. A second step aims to obtain a template by averaging the obtained features, after which the third step compares subsequent breaths to this template. This allows for the exclusion of abnormal breaths with strange artifacts, the reasons for which can be inferred by comparing the features with classes of breaths with known abnormalities. In a similar approach, breaths are separated by precisely locating positive and negative gradients, after which a template waveform for CO2 values is constructed based on the identified averages and outliers.
[0138] A rule-based approach was also proposed, which defined the criterion of absolute change in CO2 concentration as a method for extracting the phase of respiratory separation, and further established a rule including the duration of respiration for respiratory exclusion. Other methods search for the phase of respiratory separation in a sliding window. The onset and completion of a breath are identified using inflection points at the troughs of the breath, and breaths are excluded based on pre-established extremes of physiological likelihood. Another preprocessing approach involves fitting a model to the breaths. For example, the breaths are segmented into phases by fitting a series of piecewise linear lines, and then the breaths are segmented into individual breaths by applying a logistic algorithm to the troughs. Similar techniques have been used to identify segments by linearly modeling the data in a sliding window of length 3 and applying a rule to a plot of the slope of this fit over time.
[0139] Machine learning has also been used for breath segmentation, where an artificial neural network (ANN) ingests the raw time series and outputs the start and end points of a breath. A similar solution uses an ANN on eight features derived from the respiratory waveform for bad breath rejection.
[0140] These capnographs were primarily obtained from sedated patients at low sampling rates, resulting in greater homogeneity in the capnograph presentation.
[0141] None of these solutions provides an accurate, reliable technique for processing high frequency noisy signals with artifacts, resulting in interpretable and computationally efficient processing and classification.
[0142] Figure 4 A method for predicting whether a capnogram is produced by a user with cardiopulmonary disease according to an example of the present disclosure is shown. Known methods cannot handle various artifacts that may affect the automatic detection of turning points.
[0143] In the first step, the method retrieves a single-breath waveform from a capnogram (S101). This can be in the form of a digitized sample as discussed above. Next, a set of turning points can be identified (S102). In one example, these turning points are identified by analyzing the first-order differential of the waveform and the maximum and minimum values of the signal. Preferably, the first-order differential is calculated using a time-based smoothing filter in conjunction with the first-order derivative. Local minima and maxima of the signal can be caused by respiratory or physiological artifacts and can interfere with the detection of turning points. Therefore, the output of this step can be a set of starting and ending points for each phase, i.e., for example, the index of the relevant sample, and optionally the sum of the first-order differential at each identified point. From these elements, the method extracts a set of features (S103). As described below, any number of features can be extracted from a classifier capable of learning from these features to predict outcomes from an unseen set of capnogram data. In the implementation phase, a trained machine learning model can be applied to the extracted features to predict the presence of disease in the user from whom the capnogram was generated (S104).
[0144] It should be understood that in a typical machine learning process, there may be three phases: training, validation, and testing. Each of these phases may be implemented by different character entities derived from features extracted from various datasets using the above method.
[0145] In its general form, a supervised machine learning model takes as input data: m instances and n features sampled , and output data in pairs: , as an m×k matrix corresponding to m instances and k outputs. The goal is to create a function using algorithm A that uses this data , this function can get any unknown And output the prediction : .
[0146] In this problem statement, X is the set of characterized capnography waveforms, and This is a label that indicates whether the capnogram is from non-COPD or COPD-affected lungs. This distinction is made between healthy and non-COPD lungs because patients without COPD may have other conditions that affect the capnogram waveform.
[0147] In step S101, a respiratory waveform is obtained from a capnogram. When the capnogram includes multiple breaths (these breaths may be full breaths and partial breaths), Figure 2 In the example of FIG, obtaining a respiratory waveform includes dividing the capnogram into a plurality of segments. When dividing the capnogram into a plurality of segments, each capnogram segment should preferably consist of a single respiratory waveform corresponding to a single respiratory cycle (i.e., a complete breath or a partial breath if sampling of the capnogram data starts or stops mid-breath).
[0148] A particular benefit of this method is the ability to accurately and reliably make predictions based on a single waveform. Preprocessing steps such as denoising and breath separation are preferably performed, but these are not required and are not described. Another preprocessing step, which forms a stylized single waveform, is another example of the present invention and is described below.
[0149] In step S102, one or more turning points of the waveform are determined. Specifically, the determined turning points include one or more of a delta turning point δ, a gamma turning point γ, and an alpha turning point α.
[0150] These turning points can be determined by a variety of methods. Due to its high accuracy and compatibility with automation and machine learning, a preferred method involves determining the first differential of the respiratory waveform. It has been found that the first differential of the respiratory waveform can be used as a reference for the respiratory waveform itself to determine several turning points.
[0151] Determining the first-order differential of a waveform also allows for screening for respiration viability to conserve computational resources. For example, the first-order differential can be compared to a differential template, and when the first-order differential of a respiratory waveform is inconsistent with the template, the respiratory waveform is excluded and analysis of the waveform ceases. The nature of the template and the degree of inconsistency allowed before the threshold for exclusion can vary depending on the specifics of the waveform analysis; for example, depending on which turning points are to be identified and which features of the respiratory waveform are to be extracted.
[0152] 5A to 5D Several examples of respiratory waveforms overlaid by their first-order differentials are shown, where the first-order differentials are determined in conjunction with a Savitsky-Golay (SG) filter applied to the respiratory waveform. The scale of the SG filter first-order differentials and the respiratory waveforms has been changed in these figures to facilitate easier comparison of the SG filters and the respiratory waveforms. 5A to 5D It is obvious that, especially when Figure 5A and Figure 3 When compared with the idealized waveform of the respiratory waveform, the peak of the first-order differential can basically correspond to the turning point of the respiratory waveform, or be close to the turning point, and thus can be used as a part of determining the turning point.
[0153] from Figure 5B 、 Figure 5C and Figure 5D It is clear from the graph that artifacts in the respiratory waveform have a significant effect on the shape and first-order differential of the waveform. For example, Figure 1C 、 Figure 5B and Figure 5C 2 hump artifacts. These artifacts should preferably be identified and processed during waveform analysis (e.g., during determination of turning points and / or feature extraction). This is particularly beneficial when the step of determining turning points is automated.
[0154] Preferably, when a hump artifact is identified in a respiratory waveform, the hump artifact is addressed during further analysis of the respiratory waveform. For example, if the hump artifact produces a maximum amplitude point on the first-order differential that is not caused by the expiratory ascender P2 or the inspiratory descender P4a, the first-order differential can be normalized relative to the peak value caused by the ascender P2 or descender P4a (rather than the peak value caused by the hump artifact) to address the hump artifact. Various thresholds used to determine turning points (discussed below) can then be defined based on the renormalized first-order differential to avoid the influence of the hump artifact. Alternatively, the artifact can be addressed by ignoring the artifact, for example, by ignoring or removing data related to the artifact (effectively, subtracting the artifact), or by adjusting the weighting applied to the artifact-related data when determining one or more turning points.
[0155] Artifact identification can be performed by several methods. Due to high accuracy and compatibility with automation and machine learning, a preferred method includes performing a peak detection operation on the respiratory waveform and optionally also includes determining and / or using the first order differential of the respiratory waveform.
[0156] As part of artifact identification, peak detection is performed to identify local minima of the respiratory waveform. Significant minima can then be identified from the local minima. Identifying significant minima is particularly useful when the respiratory waveform is a noisy signal. When one or more significant minima are found, local regions of the waveform are examined to determine if a hump artifact is present, and if so, the hump artifact is identified. For example, from Figure 1C It can be clearly seen in FIG that there is a significant local minimum before the expiratory ascending limb P2 and after the hump artifact. Searching for the local maximum near this significant minimum will identify the hump artifact before the pre-expiratory ascending limb P2.
[0157] When no significant minimum is identified (e.g., due to a noisy respiratory waveform), the first-order differential of the respiratory waveform can be used to search for hump artifacts. Figure 5A and Figure 5B and Figure 5C As shown by the comparison of , the first-order differential (in this case, the SG filter) is significantly affected by the presence of the P2 hump artifact in the pre-expiratory period (i.e., during the expiratory baseline P1 period) and can therefore be used to identify hump artifacts that would otherwise be missed. Figure 5C In the figure, the SG filter increases, then briefly reaches a plateau at the inflection point, and then increases again, indicating that there is a smaller (relative to Figure 5B ) Camelback artifact.
[0158] When a hump artifact is identified, whether or not the first-order differential is used, the hump artifact must be processed during the turning point determination process to ensure accurate definition of the turning point. For example, after the hump artifact is identified, the first-order differential of the respiratory waveform can be normalized or renormalized based on the maximum amplitude minimum / maximum point of the first-order differential excluding the hump artifact.
[0159] In order to avoid unnecessary use of computing resources, the region of the respiratory waveform considered during these steps (e.g., peak detection and / or significant minimum identification) can be limited. For example, the maximum value of the respiratory waveform is identified, and the respiratory waveform is temporally divided into a first portion and a second portion. The first portion is a time period of the respiratory waveform that does not include the maximum value of the respiratory waveform, and the second portion is a time period of the respiratory waveform that includes the maximum value of the respiratory waveform. In order to save computing resources and avoid searching for hump artifacts during the inspiratory phase (P4a and P4b), the steps for identifying hump artifacts (e.g., identifying local minima, or searching for hump artifacts when at least one significant minimum is identified) are performed in the first portion of the respiratory waveform rather than in the second portion of the respiratory waveform.
[0160] It has been found that minima corresponding to hump artifacts in the capnography waveform typically have pCO2 values below 2 kPa. Given this, a preset threshold (herein, the hump artifact threshold) can be implemented to further reduce the processing resources required when searching for hump artifacts. This can be achieved in a variety of ways, such as by discarding any detected minima above the hump artifact threshold, or by not searching for minima above the hump artifact threshold in the first example. The hump artifact threshold can be used in place of the first and second partial partitioning described above, or in combination with the hump artifact partitioning of the respiratory waveform.
[0161] Even if no significant minimum is identified, hump artifacts may still be present in the respiratory waveform. This could be, for example, Figure 5CThis depends on the peak detection method used or the technique used to identify significant minima. When no significant minima are identified, the first-order differential of the waveform can be used to identify whether the respiratory waveform contains a hump artifact. The first-order differential is examined for the inflection point region, and the differential in the inflection point region can be compared to a preset threshold.
[0162] Preferably, in order to reduce the processing resources used, only the first-order differential in the region between the start point of the first-order differential and the maximum value of the first-order differential is analyzed.
[0163] Alternatively, local maxima (and corresponding significant maxima) of the respiratory waveform may be used to identify hump artifacts, rather than identifying local minima of the respiratory waveform and significant minima from among the local minima.
[0164] When hump artifacts are present, steps should be performed to identify and process any hump artifacts before identifying turning points to most accurately define turning points.
[0165] The first-order differential of the respiratory waveform can be used to determine the delta turning point δ. A preferred method involves determining the maximum amplitude point of the first-order differential (i.e., the value of the highest peak or lowest trough of the first-order differential) and using this maximum amplitude point to define a preset threshold (referred to herein as the delta threshold). For example, the delta threshold can be a value equal to 10% of the maximum amplitude point of the first-order differential. This delta threshold (and other thresholds) can be found by normalizing the first-order differential to 1 relative to the maximum amplitude point of the first-order differential (after processing any hump artifacts). The location of the delta turning point δ in the time dimension can then be defined as the first time point at which the first-order differential of the respiratory waveform is above the delta threshold (after processing any identified hump artifacts). That is, the respiratory waveform value corresponding to this time point is the delta turning point δ between the expiratory baseline P1 and the expiratory ascending limb P2.
[0166] The gamma turning point γ can also be determined using the first-order differential of the respiratory waveform. The preferred method includes determining the maximum amplitude point of the first-order differential (i.e., the value of the highest peak or lowest trough of the first-order differential), and using the maximum amplitude point to define a preset threshold (referred to herein as the gamma threshold). For example, the gamma threshold can be a value equal to minus 5% of the maximum amplitude point of the first-order differential. The method also includes determining the minimum value of the first-order differential. The position of the gamma turning point γ in the time dimension can then be defined as the first time point after the minimum value of the first-order differential of the waveform, at which time point the first-order differential is above the gamma threshold. That is, the respiratory waveform value corresponding to this time point is the gamma turning point γ between the inspiratory descending limb P4a and the inspiratory baseline P4b.
[0167] The alpha turning point α can also be determined using the first-order differential of the respiratory waveform. A preferred method includes determining the maximum amplitude point of the first-order differential (i.e., the value of the highest peak or lowest trough of the first-order differential) and using this maximum amplitude point to define a preset threshold (referred to herein as the alpha threshold). For example, the alpha threshold can be a value equal to 15% of the maximum amplitude point of the first-order differential. The method also includes determining the maximum value of the first-order differential and the maximum value of the respiratory waveform. In some examples, the maximum amplitude point of the first-order differential can also be the maximum value of the respiratory waveform, but this is not always the case and will depend on the specific respiratory waveform being analyzed. The position of the alpha turning point α in the time dimension can then be defined as the first time point after the maximum value of the first-order differential, at which time point the first-order differential is less than the alpha threshold. That is, the respiratory waveform value corresponding to this time point is the alpha turning point α between the expiratory ascending limb P2 and the expiratory plateau P3.
[0168] In some examples of this method, the alpha threshold can be adjusted. For example, when there are no points between the maximum value of the first-order differential and the maximum value of the respiratory waveform where the first-order differential is less than the alpha threshold, the alpha threshold can be increased, and the value of the first-order differential can be compared to the increased alpha threshold (between the same boundaries). This process can be performed iteratively until an alpha turning point is defined, or until an upper limit on the number of iterations performed is reached. If this limit is reached, the respiratory waveform can be excluded, or analysis of the respiratory waveform can continue.
[0169] To conserve processing resources, the location of the alpha turning point α in the time dimension can be defined as the first time point between the maximum of the first-order differential of the respiratory waveform and the maximum of the respiratory waveform (or the beta turning point β) at which the first-order differential is less than the alpha threshold. Since the alpha turning point α is known to be located in the respiratory waveform before the beta turning point β and is likely to be before the maximum of the respiratory waveform, these restrictions prevent unnecessary data analysis when determining the alpha turning point α.
[0170] An alternative method for determining the alpha turning point α includes calculating a linear line between the delta turning point δ and the maximum value of the respiratory waveform (or the beta turning point), and defining the alpha turning point α based on the distance between the respiratory waveform and the calculated line. For example, when the distance between the respiratory waveform and the line is measured using a second linear line perpendicular to the calculated line, the alpha turning point α can be defined as the point of the respiratory waveform between the delta turning point δ and the maximum value of the respiratory waveform (i.e., the point farthest from the calculated line).
[0171] Determining one or more turning points may also include determining a beta turning point β between the expiratory plateau P3 and the inspiratory descending limb P4a. The beta turning point β can be used to extract additional features from the respiratory waveform and help determine or verify other aspects of the waveform, such as the alpha turning point α and the hump artifact.
[0172] A peak detection operation can be used to determine the beta turning point β. The peak detection operation is performed to identify local maxima of the respiratory waveform. Significant maxima can then be identified from the local maxima; this is particularly useful when the respiratory waveform is a noisy signal. Significant maxima can be identified by comparing each identified maximum to a threshold or by comparing within the set of maxima. Optionally, because the CO2 values at the beta turning point β are expected to be above a preset threshold (herein, the maximum threshold) for a valid respiratory waveform, CO2 values below the maximum threshold can be ignored (e.g., by removing these values during peak detection or setting them equal to 0) to reduce noise from values below the maximum threshold during peak detection.
[0173] When no significant maximum is identified, this typically indicates that the respiratory waveform is a poor sample and, therefore, the waveform is excluded without further unnecessary processing. When only a single significant maximum is identified, the location of the beta turning point β in the time dimension can be defined as the time point of the single significant maximum. When multiple significant maxima are identified, the most significant maximum is determined and considered the beta turning point β. A variety of methods can be used to determine the most significant maximum; one method uses a two-step significance-based algorithm that compares the height of each maximum to the surrounding area, and if two or more maxima are not filtered out at this stage, the heights of the remaining maxima are compared. If, for example, the most significant maximum cannot be determined within a reasonable number of iterations, the first of these significant maxima is considered the beta turning point β. If other turning points are known, these points can be used to help determine the beta turning point β. For example, the beta turning point β precedes the gamma turning point γ, so any abnormally significant maxima after the gamma turning point γ are not considered during the determination of the beta turning point β.
[0174] When the beta turning point β is known, it can be used in place of the maximum value of the respiratory waveform in the methods described above. That is, the position of the alpha turning point α in the time dimension can then be defined as the first time point between the maximum value of the first-order differential of the respiratory waveform and the beta turning point β at which the first-order differential is less than the alpha threshold.
[0175] The beta turning point β can also be used in an alternative method for determining the alpha turning point α described above. Instead of calculating a linear line between the delta turning point δ and the maximum value of the respiratory waveform, a linear line is calculated between the delta turning point δ and the beta turning point β. The alpha turning point α can then be determined based on this calculated line in the same manner as described above.
[0176] Similarly, when identifying a hump artifact using the method described above, the respiratory waveform can be temporally divided into a first portion that does not include the beta turning point β and a second portion that includes the beta turning point β. As described above, the steps for identifying a hump artifact are performed in the first portion of the respiratory waveform, rather than in the second portion of the respiratory waveform (e.g., identifying a local minimum or searching for a hump artifact when at least one significant minimum is identified).
[0177] We have already described how to use first-order differentials (or higher-order differentials or a combination thereof) to identify turning points and algorithmically compensate for noise and artifacts. Another approach could be to analyze point-by-point differences in the sample. For example, this approach would take point-by-point differences in the sequence, i.e., if the sequence is {t0, t1, t2, …, t n}, this will produce {t1-t0, t2-t1, …,t n -t n-1}, which is 1 shorter than the original sequence. t0 will be added to the starting point, so the final sequence is {t0, t1-t0, t2-t1, …, t n -t n-1}.
[0178] That is, the difference between each value or sample represents the change in the capnography curve. This can be analyzed to identify turning points, dealing with noise and artifacts in a similar manner as described above. In summary, in addition to the described methods with time-based smoothing filters and first-order derivatives, higher-order derivatives or point-by-point differences in time-series capnography data (i.e., the values of the curve) can also be used to identify the shape of the curve and turning points between phases.
[0179] Figure 6A 、 Figure 6B and Figure 6C These above steps are shown in the form of a flow chart.
[0180] exist Figure 6AIn the first step shown, the inspiration and expiration phases are extracted by identifying the starting position of the inspiration phase. The input to this process is a digitized sample of capnography data 601. A first-order Savitzky-Golay differential is calculated using an appropriate Savitzky-Golay filter 602 and normalization 603. If this is inconsistent with good breathing, it can be rejected, for example, by comparison with a template shape of the differential, or by threshold comparison or other suitable means of identifying the differential as desired.
[0181] Next, a peak detection algorithm is used to find local maxima and minima 604 in the sampled data 601. Based on the subset of minima, if no significant minima are detected (e.g., identified using a bidirectional local window 605), the differential is analyzed 606 to identify additional humps before the onset of exhalation. If these exist 608, the differential can be renormalized 603 and the process begins again. The differential can identify the absence of humps in the sampled data. If multiple minima are found in the subset, the samples are checked for minima in the first half of the breath (i.e., cycle) where the minimum is less than a threshold (i.e., hump artifact threshold). The threshold can preferably be selected to be less than 2 kPa, which is selected to identify humps that are unlikely to be on a plateau.
[0182] Based on the local maxima identified by the peak detection algorithm, it can be identified that the start of inspiration is at the end of the breath. Such edge cases can be handled appropriately. Based on a subset of the most significant maxima 610 (e.g., identified by a one-way positive window), if only one maximum is detected, the location of the start of inspiration 611 can be identified. If no maximum is found or if the resulting phase is too short, the waveform can be rejected. If more than one significant maximum is identified 612, the iterations described above can be performed to identify the start of inspiration by comparing each maximum to a threshold or to each other, or alternatively, the first maximum can be taken as the start of the inspiratory phase.
[0183] exist Figure 6B In the second step shown, based on the normalized differential, i.e., the SG filter differential, and the position in the waveform defined as the start of inspiration, the method can identify the breathing turning points and the start and end of each phase of breathing.
[0184] The location 613 at the end of Phase 1 can be identified as the first location where the differential exceeds a positive threshold 614. Preferably, the threshold can be 0.1, when the differential has been normalized to 1 relative to its maximum amplitude point (after processing any hump artifacts). The threshold is positively selected through careful study and analysis. If the differential does not exceed the threshold, the waveform can be rejected.
[0185] Based on this differential, the differential between the inspiration start location and the maximum point of the differential can be compared to a threshold 615. If many points are identified, i.e., more than 20 locations, the waveform can be excluded. The threshold can be iteratively incremented 616, for example, by 0.05, until fewer points are identified. The location of the end of Phase 2 can then be identified 617 based on this comparison 615. Based on the two identified locations 613, 617, the duration of Phase 2 can be identified 618.
[0186] The end of phase 3 can be identified based on the previously identified onset of inspiration 611. Thus, the duration 619 of phase 3 can be identified by comparing the end of phase 2 617 to the onset of inspiration 611.
[0187] Based on the differential, the first point 620 that exceeds the negative threshold between the differential minimum and the differential end point can be identified. Thus, the location of this point and the location of the differential minimum can be used to identify the end point 621 of phase 4a. The end point 622 of phase 4b can be defined as the range between the end point of phase 4a and the end point of the waveform.
[0188] The duration 612 of phase 4a may be a range defined as the position of the start point 611 of the aforementioned inhalation phase (ie, the end point of phase 3) and the end point 621 of phase 4a.
[0189] therefore, Figure 6A and 6B The outputs of the two illustrated processes in may be the normalized first differential 603, the inspiration start position 611 and the start and end points of each phase of the waveform, ie the duration or extent 613, 618, 619, 612, 622 of the phase.
[0190] In another optional example of the present disclosure, the segmentation can be further divided into angular phases. An example of such a process flow chart is shown in Figure 6C Shown in.
[0191] The start 623 of the delta angle is defined as the end 614 of phase 1. The end 626 of the delta angle is defined as the first significant maximum 624 of the differential 603, which reaches its minimum 626 after the local minimum 625 but before the last point of the differential. The local minimum 625 is used to avoid any hump at the beginning of the waveform.
[0192] The start of the α angle 627 is defined as the first time after the first significant maximum 624 that the differential falls below a threshold 628, which is preferably selected within the range of 0.2 to 0.9, and more preferably selected to be 0.5. The end of the α angle 629 is defined as the end of phase 2.
[0193] The start of the angle β is defined as the inspiration start position. The end of the angle β 632 is identified as the first time after inspiration start that the differential is less than a negative threshold 631, which is preferably selected within the range of -0.4 to -1, and more preferably selected as -0.9.
[0194] The start of the γ angle is defined as the first time after the differential minimum 625 that the differential exceeds a negative threshold 634, preferably selected within the range of -0.6 to -1, and more preferably selected to be -0.9. The end of the γ angle may be defined as the end 621 of stage 4a.
[0195] These angular features may optionally, and particularly preferably, be used to identify further beneficial features based on turning points, as discussed below.
[0196] In step S103 , the determined turning points are used to extract features of the respiratory waveform.
[0197] Deriving important features (i.e., biomarkers) from the capnography waveform allows for training and configuring machine learning models to make predictions based on these features. Many different features of the respiratory waveform can be extracted using various methods; a non-exhaustive list of examples is provided below. It should be noted that the benefits of each of these identified features can be separated from the overall approach in which we present them.
[0198] Extracting biomarkers or features from capnography has become a significant area of research. Each of these features is particularly beneficial, and their description and automated calculation may not have been previously reported. In any case, extracting such features based on turning points has not been previously described, nor has the efficient calculation of these features without human intervention been described.
[0199] In the first set of features, the angle at the turning point can be determined by applying a linear fit to each phase on either side of the turning point and calculating the angle between them. For example, the alpha angle (the angle at the alpha turning point α) can be determined by applying a linear fit to the expiratory ascending limb P2 and the expiratory plateau P3 and calculating the angle between them at the alpha turning point α where the fitted line intersects. A linear line can be fitted between two turning points (e.g., a linear line fitted to the expiratory ascending limb P2 between the delta turning point δ and the alpha turning point α), or between the target turning point (in this case, the alpha turning point α) and another point along an adjacent phase (e.g., a point along the expiratory ascending limb P2 or the expiratory plateau P3 in this case). The nature of the linear fit can depend on the shape of the respiratory waveform. It has been found that different features indicate different physical properties. For example, in the context of the features and turning points defined herein, the alpha angle indicates obstructive airway disease.
[0200] As mentioned above about Figure 6C As described, the angle start and end values (e.g., time values and CO2 values) are also important features that have been found to be consistently effective in driving machine learning models and improving their performance, and the angle start and end values also lead to further features that improve machine learning models. Other examples of these features are described below.
[0201] The delta angle start can be defined as the delta turning point δ (i.e., the end of the expiratory baseline P1), while the delta angle end requires further calculation to determine. The delta angle end can be determined using the first-order differential of the respiratory waveform. The delta angle end position in the time dimension can then be defined as the time point corresponding to the first significant maximum of the first-order differential of the respiratory waveform (from the start of the respiratory waveform, or, when identified, after the hump artifact). Alternatively, the delta angle end can be defined as the first significant maximum (i.e., relative to a preset threshold) of the differential after a local minimum of the differential and before the minimum.
[0202] The position of the α angle start point in the time dimension can be defined as the first time point after the first significant maximum of the first-order differential (i.e., the first point after the δ angle end point), at which the first-order differential of the respiratory waveform falls below a preset threshold (referred to herein as the α start threshold). For example, the α start threshold can be a value between 20% and 90% of the maximum amplitude point of the first-order differential, and more preferably, 50% of the maximum amplitude point of the first-order differential. The α angle end point can be defined as the α turning point α (i.e., the end of the expiratory ascending limb P2).
[0203] The β angle start point can be defined as the beta turning point β (i.e., the start point of the inspiratory descending limb P4a). The position of the β angle end point in the time dimension can be defined as the first time point after the β angle start point at which the first-order differential of the respiratory waveform falls below a preset threshold (referred to herein as the β end point threshold). For example, the β end point threshold can be a value between -40% and -100% of the maximum amplitude point of the first-order differential of the respiratory waveform, and more preferably, a value of -90% of the maximum amplitude point of the first-order differential of the respiratory waveform.
[0204] The position of the γ angle start point in the time dimension can be defined as the first time point after the minimum value of the first-order differential of the respiratory waveform, at which the first-order differential exceeds a preset threshold (referred to herein as the γ start point threshold). For example, the γ start point threshold can be a value of minus 60% of the maximum amplitude point of the first-order differential of the respiratory waveform, and more preferably, a value of minus 90% of the maximum amplitude point of the first-order differential of the respiratory waveform. The γ angle end point can be defined as the gamma turning point γ (i.e., the end of the inspiratory descending limb P4a).
[0205] In certain embodiments, because there are often small variations in the number of breaths in each capnogram and in the shape of breaths within a capnogram, a median can be calculated for each breath feature in each capnogram. Example features in certain embodiments include the following: alpha, beta, gamma, and delta angles; gradients and residuals from curve fits to phases (e.g., expiratory plateaus); absolute and short-term variability in pCO2; curvature and other higher-order time-based features, such as the ratio of expiratory to inspiratory phases; second-order area ratios and area under the curve (AUC) in quadrants of the expiratory phase, such as those calculated in volumetric capnograms.
[0206] Features associated with the alpha region of the capnography waveform have been found to be the most important driver of learning. The discriminative features that contribute most to the model's decision making have been found in the alpha angle region, which characterizes the rate at which gas from the upper airway (low CO2) is replaced by mixed alveolar gas from the lower airway (high CO2).
[0207] In step S104, the trained machine learning model is applied to the extracted features. The trained machine learning model is configured to predict the presence of cardiopulmonary disease based on the features. In particular, the trained machine learning model is configured to predict whether the capnogram (from which the features were extracted) was generated by a patient with cardiopulmonary disease.
[0208] Examples of applicable machine learning algorithms include logistic regression, gradient boosting decision trees, support vector machines, AdaBoost, and random forests, however, it will be appreciated that other algorithms are also suitable. Some machine learning models, such as logistic regression models, allow analysis of the impact of each extracted feature on the prediction. Some models, such as gradient boosting decision tree models, maintain a high level of interpretability while adapting to overfitting, helping to ensure that the model can be generalized to new data. It will be appreciated that the present invention can be implemented using various machine learning models other than those listed in the specification.
[0209] A trained machine learning model can be configured to provide additional information as part of or alongside a prediction. For example, a trained machine learning model can predict the type of cardiopulmonary disease, or the severity or subtype of cardiopulmonary disease, that a user has when generating a capnography waveform. The model can also provide a level of certainty associated with the prediction and / or highlight the specific features that led to the prediction and the degree to which (different) features contributed.
[0210] As part of step S104, the trained machine learning model can be applied to features extracted from a single respiratory waveform or from multiple respiratory waveforms. The multiple respiratory waveforms can be from a single capnogram (i.e., a series of respiratory waveforms recorded by the user during a single respiratory period) or from multiple capnograms taken at different times. When extracting features from multiple capnograms, these capnograms can be obtained (i.e., recorded by the user) over an extended period of time. For example, capnograms can be recorded repeatedly by the user multiple times a day over a period of days, weeks, or months. Extracting features from multiple respiratory waveforms obtained over an extended period of time (i.e., extracting features from respiratory waveforms in multiple capnograms rather than from a single capnogram) allows for determining the variability of the extracted features and provides a more accurate prediction of whether the user suffered from a cardiorespiratory condition at the time their capnogram was generated. For example, it has been found that the cardiorespiratory condition asthma can be more accurately predicted when the variability of the extracted features is assessed using at least three capnogram recordings, preferably those recorded over a period of at least five days, and more preferably those recorded over a period of at least ten days. In contrast, the variability of the extracted features may be less important for predicting the presence of the cardiopulmonary disease COPD, and thus COPD can be predicted very accurately based on a single respiratory waveform using the disclosed method.
[0211] Figure 7 A method for training a machine learning model to learn metrics from capnograms produced by users with cardiopulmonary diseases is shown.
[0212] In step S201, a respiratory waveform is obtained from each of the plurality of capnograms. Step S201 mirrors step S101 by segmenting the capnogram into portions consisting of a single complete respiratory waveform. Training a machine learning model using a larger amount of data with a wider variety of characteristics generally results in a better trained model capable of making more accurate predictions. Preferably, the larger amount of data is provided by obtaining respiratory waveforms from a greater number of different capnograms (i.e., most preferably, different capnograms generated by different airways at different times). However, multiple respiratory waveforms can also be obtained from a single capnogram in the plurality of capnograms.
[0213] In step S202, one or more turning points of the waveform are determined. Step S202 corresponds to step S102 described above and can be performed using the same techniques. This also includes, for example, determining the first-order differential of the respiratory waveform and identifying and processing hump artifacts.
[0214] In step S203, the one or more turning points determined in step S202 are used to extract features of the respiratory waveform. Step S203 corresponds to step S103 described above and may be performed using the same techniques.
[0215] When steps S202 and S203 are repeated for each of the different respiratory waveforms in the plurality of capnograms, it is not necessary to determine the same turning point for each respiratory waveform and / or extract the same features from each waveform. However, it is preferable to determine the same turning point and extract the same features for multiple respiratory waveforms (and more preferably, for most respiratory waveforms) in order to more efficiently train the machine learning model.
[0216] In step S204, a label is obtained for each capnogram, where the label indicates whether the capnogram is from a diseased airway (or a healthy airway). Preferably, when the label indicates that the capnogram is from a diseased airway, it also indicates the presence of a cardiopulmonary disease affecting the airway at the time the capnogram was recorded. For example, when training a machine learning model to learn metrics from capnograms generated by users with COPD, it is preferable to label the capnograms as "COPD" or "non-COPD" rather than "COPD" or "healthy," as airways without COPD may harbor other cardiopulmonary diseases that affect the waveform of the generated capnogram.
[0217] Obviously, step S204 can be executed before or after any one of steps S201, S202 and S203.
[0218] In step S205, the labels and extracted features of the plurality of capnograms are used to train a machine learning model to learn indicators of cardiopulmonary disease to create a prediction function. The resulting prediction function is a trained machine learning model suitable for use in step S104.
[0219] In a preferred example, Figure 8 The schematic diagram shows the combination of multiple machine learning methods. Figures 9A to 9C Shown Figure 8 The three parts.
[0220] Based on a representation of a single breath waveform 801, including, for example, an average waveform or a template waveform, feature engineering 802 is performed as described below. Multiple machine learning models are then used in three example components 803, 804, 805 of the system.
[0221] Three different ML models were selected in this example: logistic regression (LR), gradient boosted decision tree (XGBoost), and support vector machine (SVM). LR and XGBoost were chosen as simple, effective, and interpretable algorithms, while SVM, more specifically kernel SVM, was included to perform classification on high-dimensional, nonlinear feature spaces. Deep learning methods could also be used to maintain interpretability and avoid increased computational complexity, but are not preferred in these examples. It should be understood that deep learning methods such as ANNs or BiLSTMs are also applicable to the methods of the present invention, and are particularly suitable for feature engineering.
[0222] Figure 9A A schematic diagram illustrates an embodiment for identifying COPD patients based on a single representation of their respiratory recordings. Patients 901 are partitioned into training and validation datasets 902, and a model is trained 903 on 902. Optionally, capnograms are obtained and grouped according to patient-stratified k-fold cross-validation, where each patient's recording is present only within a single fold of the partition. Another subset 904 can optionally be used as an unknown test set input to the trained model 905, which is then learned to classify COPD 906 or non-COPD 907.
[0223] Most existing methods focus on classification accuracy without providing specific / interpretable information that would be useful to clinicians or experts. Furthermore, other products only provide basic information, such as end-tidal CO2 and respiratory rate, which do not provide sufficient information to distinguish between various diseases. Existing methods are generally designed to classify COPD versus normal; if examined, patients are most likely to show symptoms of lung disease, and therefore are more likely to have other diseases, so it is preferable to train models on COPD versus non-COPD versus normal / healthy.
[0224] Figure 9B It shows how to extend the classification task to a multi-class classifier, for example a three-class classifier. The three classes are COPD 909, asthma 908, and healthy 910. The same method and features are used in the binary task of COPD vs. non-COPD.
[0225] Figure 9C A further extension is shown, where the task can be extended using time series methods to analyze how features change over time.
[0226] Greater heterogeneity has been observed in patients with asthma, with some patients with severe asthma observed to have capnography waveform geometry closer to that of COPD, while others are closer to healthy individuals. This has not been a particularly problematic issue before, as it is difficult to obtain longitudinal capnography data from multiple different patients, and alternative sensors are not accurate or fast enough to effectively address rapid transitions. Trying to identify patients with asthma based on a single metric would be very challenging. The present method is able to capture this longitudinal information.
[0227] Longitudinal information can include multiple respiratory records from the same acquisition period, or multiple capnograms acquired over time. A time series approach is introduced to calculate the variability over an n-day window, where the variability is defined as the standard deviation between all features extracted. Thus, the third system classifies between healthy and asthmatic patients. Figure 9C As shown, data is collected from a patient over a period of time. As indicated above, in an example, the features extracted from each waveform at a time can be compared, for example, to identify the mean, median, or standard deviation of the features, or alternatively, to identify the statistical distribution of the features over time. These can be used to further train a model based on the variability of the features.
[0228] like Figure 9C As shown, patient 901 undergoes a time series method 911 before being divided into training and validation datasets and an unknown test dataset. When training model 905 is trained 903, the patient data is classified into asthma and healthy.
[0229] The algorithmic method S102 for identifying turning points of a capnogram has been described above. In an alternative embodiment, a machine learning method may be used to identify turning points (S102).
[0230] To train the model, multiple respiratory waveforms may be obtained. As described above, the respiration may be in the form of a set of discrete samples. A set of labels may be applied to one or more samples of each respiratory waveform. The labels may represent features that the machine learning model is intended to learn. For example, a label may correspond to the phase of a particular respiratory waveform. Alternatively, a label may correspond to a positive or negative indication of whether the sample corresponds to a turning point. Alternatively, a label may correspond to a real number indicating the distance in the respiration at which a particular turning point occurs. In a preferred example, there may be five output classes, with each label identifying to which of the five classes the sample belongs.
[0231] The set of labels and training respiratory waveforms can each be provided to a classifier or other suitable machine learning model, such as a convolutional neural network, which can be configured to learn to predict turning points for an unknown set of samples. In a specific example, a logistic regression model can be trained in a supervised manner based on the defined labels and samples to learn to classify each individual sample into a defined set of output classes. The input can be an entire breath normalized to a fixed length. Similarly, there can be multiple logistic regression models, one for each sample.
[0232] In the implementation phase, a set of unknown samples of respiration can be fed to a model (or, in the example above, multiple models, where a trained model exists for each sample) trained based on the aforementioned labels and training respiration waveforms. The unknown samples are classified into defined output classes. Based on the samples adjacent to the boundaries of each class, turning points between phases of the unknown waveform can be identified. This information can then be used to extract features of the waveform, as identified elsewhere in this disclosure.
[0233] Disease classification
[0234] The preceding article describes how to train a classifier or other machine learning model to identify the likelihood of a cardiopulmonary disease being associated with an unknown respiratory waveform or capnography pattern. It should be understood that while the concepts here are described using supervised learning terminology, they are equally applicable to unsupervised learning. While classification has been described as an example, it should be understood that clustering techniques can also be used.
[0235] In a specific example, higher-complexity deep learning models can detect features in CO2 waveforms, in addition to more interpretable methods that rely on feature engineering. In this specific example, a 1D convolutional neural network can be trained on standardized CO2 waveforms. Saliency maps, a method that facilitates interpretability, can be applied to identify regions of the respiratory waveform that most contribute to predicting the likelihood of COPD.
[0236] Nonetheless, supervised machine learning and interpretable methods have important advantages in the medical field.
[0237] The proposed model can be trained and configured to identify the likelihood of capnography patterns associated with a disease or disease subtype. In a specific example, the disease is a cardiopulmonary disease, specifically COPD, asthma, asthma-COPD overlap syndrome (ACOS), or a subtype of cardiopulmonary disease, such as asthma with small airway disease. In the following description, we generally refer to diseases, but we also mean subtypes. For example, when we refer to disease A and disease B, we also mean disease A and subtype B, or subtype A and subtype B, and so on.
[0238] To predict the likelihood of disease based on a capnography waveform, a machine learning model can be configured as a classifier, where the model is trained based on labels indicating whether a respiratory waveform or capnography waveform is clinically associated with a disease or disease subtype, or whether it is not clinically associated with a disease or disease subtype, as the case may be. In this example, the model can be configured as a binary classifier trained to predict a value in the range [0, 1], where 0 corresponds to the absence of disease and 1 corresponds to the certainty of the presence of disease. In other words, each label belongs to the class {disease, non-disease}.
[0239] Based on this likelihood, a hierarchical model is proposed to categorize the risk of the disease. It is expected that a set of thresholds (e.g., 4) will be applied to stratify the likelihood into a set of categories (e.g., 5) to aid clinical rule-in / rule-out and allow the integration of risk-based approaches into clinical pathways.
[0240] In another instance, the model may not be configured as a binary classifier for disease or not, but may be configured to identify the likelihood of different diseases. Here, the machine learning model may be configured to identify a classification problem of prediction between disease A and disease B. The model can be trained based on labels, and the labels belong to categories indicating that the obtained data is associated with the corresponding disease. In this example, the model can be configured as a binary classifier trained to predict values in the range [0,1], where 0 corresponds to disease A and 1 corresponds to disease B. In other words, each label belongs to the category {Disease A, Disease B}. Similar to the above, these likelihood (or probability) values can be stratified to identify the risk of each disease.
[0241] A specific example of this approach is given in the context of asthma and asthma-COPD overlap syndrome (ACOS). Note that time series data can be used to identify disease states, and time series-based features can be used for training. In this example, labels belonging to the asthma and ACOS categories are used, and a binary classifier is trained to output the probability of either asthma or ACOS.
[0242] According to the Dutch hypothesis, asthma and airway hyperresponsiveness predispose patients to COPD. This is supported by studies showing that asthma carries a 10-fold higher risk of developing chronic bronchitis and a 17-fold higher risk of developing emphysema compared with non-asthmatics. Early and accurate diagnosis of ACOS in asthma is crucial for the introduction of COPD medications, such as long-acting muscarinic receptor antagonists, which can alleviate symptoms and slow irreversible airway remodeling.
[0243] Spirometry is currently the gold standard for diagnosing ACOS, but it is technique-dependent, nonspecific, and requires administration by a trained healthcare professional. This, along with the lack of accurate clinical criteria for ACOS, leads to significant underdiagnosis and misdiagnosis. Therefore, a simple, safe, reliable, and accurate diagnostic test is needed to identify when asthma progresses to ACOS.
[0244] In this specific example, a model (e.g., a logistic regression model) can be trained based on features derived from capnography waveforms of asthma and ACOS participants from a clinical study. Performance can be measured on an unknown test set of asthma and ACOS participants. A possible clinical application for this model is to diagnose ACOS in or out of asthma patients with the highest model confidence. The output can be stratified into high-likely asthma; possible asthma; uncertain; possible ACOS; or high-likely ACOS based on a threshold value of the model output probability value. In this example, the waveform features driving the classification are related to the alpha angle area, which indicates a shift of alveolar gas to larger airways. These features correlate with FEV1 / FVC, supporting their hypothesized role as markers of increased obstruction.
[0245] In another specific example of this concept, ACOS is detected within the context of a binary classifier for [COPD; Asthma]. In this example, the model can be trained based on the features of the N-tidal capnography waveform as described above, along with additional clinical features. To elaborate, a machine learning model can be trained, for example, based on labels belonging to the COPD and asthma categories (i.e., {asthma, COPD}), where the model is trained and configured to predict the likelihood of a capnography waveform being associated with both COPD and asthma, and output a prediction corresponding to an overlap between the two conditions, thereby indicating a risk of both conditions.
[0246] It is also contemplated to extend the machine learning problem to multi-class classification, where labels belong to multiple categories. Each category may correspond to a corresponding disease or the absence of a disease. This extension helps predict the likelihood of one of these specified categories. In one specific example, the categories may include COPD, asthma, or ACOS (or subtypes, such as asthma, the chronic bronchitis subtype of COPD, or the emphysema subtype of COPD), allowing clinicians to use the likelihood value output to assist in diagnosing three similar and overlapping results, reducing the risk of misdiagnosis.
[0247] The implication of the above description is that each method is valid using single-label classification—that is, each capnogram can be assigned one (and only one) discrete label. However, this may not necessarily be the case. For example, the problem can be extended to multi-label classification, such as {asthma, non-COPD} or other similar multi-label problems, depending on the clinical input data obtained along with the capnogram. Similarly, multi-label classification can also predict non-disease associations alongside disease associations.
[0248] Each of these techniques can be applied to the same capnogram or respiratory waveform and the outputs presented individually, or combined or weighted, to produce a clinical indication.
[0249] In a further example of the technology disclosed herein, it can be applied to quantify small airway obstruction in asthma using rapid-response capnometry and interpretable machine learning. Asthma is a major noncommunicable disease that inflames the conducting zones of the bronchial tree and is the most common chronic disease among children. In 2019, it affected an estimated 262 million people, resulting in 455,000 deaths and totaling $50 billion in direct treatment costs annually in the United States alone. Small airway disease is a little-known contributor to severe asthma, targeting the eighth and higher branches of the tracheobronchial tree, which lack cartilage support. It has been shown to triple the odds of systemic corticosteroid use and increase the odds of acute exacerbations by sixfold. The current standard measure of small airway obstruction is an abnormally low forced expiratory flow rate (% FEF 25-75%) between 25% and 75% of predicted vital capacity, as measured by spirometry. Using the proposed technique, capnometry can be used to assess the extent of small airway disease as measured by % pred. FEF 25-75%.
[0250] Capnography can be generated from a study recruiting participants with asthma. Participants can be asked to breathe normally into an N-Tidal capnography device for at least 75 seconds at a time. These capnography signals can then be denoised, separated into individual respiratory waveforms, and characterized using geometric features for each capnography waveform. An interpretable machine learning classifier (XGBoost) can then be trained on the features of the capnography waveforms from a percentage of patients (e.g., 82%) to distinguish between % pred. FEF 25-75% < 50% and ≥ 50%, and tested on the capnography waveforms from each of the remaining patients.
[0251] Severity estimation
[0252] We have described above how respiratory waveforms can be classified into a set of output categories. Additionally or alternatively, this classification can include a measure of the severity of a respiratory disease (e.g., COPD), referred to herein as a severity metric or severity index. In other words, a severity index can be calculated as part of outputting a classification. However, within this terminology, we do not intend to output a specific classification; rather, the model can be configured to output a severity index, which is within the meaning of the phrase "outputting a classification."
[0253] For the classification of respiratory diseases such as COPD, machine learning algorithms can be trained and evaluated on COPD and non-COPD participants (including COPD patients) ranging from a set of differently defined severity levels. That is, each respiratory waveform can have a defined severity, and waveforms can be grouped according to their severity. Severity estimation can be accomplished by training a separate model (e.g., logistic regression) to distinguish mild COPD from very severe COPD, where the output probability is interpreted as disease severity. The model can be evaluated on unknown test patients using five-fold nested cross-validation.
[0254] Note that although COPD is used as an example here, the proposed technique is equally applicable to other diseases. For example, the technique can be used for asthma and GINA scores, or similarly for other diseases.
[0255] COPD severity can be derived from the % predicted FEV1 in the subset of COPD patients for whom spirometry is available.
[0256] In pulmonary function testing, a post-bronchodilator FEV1 / FVC ratio < 0.7 is generally considered diagnostic of COPD. The Global Initiative for Chronic Obstructive Lung Disease (GOLD) system categorizes airflow limitation into multiple stages. Among patients with FEV1 / FVC < 0.7, GOLD 1 (mild): FEV1 ≥ 80% of predicted; GOLD 2 (moderate): 50% ≤ FEV1 < 80% of predicted; GOLD 3 (severe): 30% ≤ FEV1 < 50% of predicted; and GOLD 4 (very severe): FEV1 < 30% of predicted. When we use the term "mild," we can refer to GOLD category 1, while severe can refer to GOLD category 4. Therefore, the proposed technique can be used to predict the classification of unknown waveforms into these categories, or to quantify severity using a metric. Similarly, respiratory waveforms assigned to these categories can be used to train machine learning models.
[0257] In a specific embodiment example, a nested cross-validation scheme can be used for model training and evaluation. Using a group-stratified cross-validation process, the full training data set can first be split into five "outer loop" folds, ensuring that a single patient's capnogram is not split between the outer loop training set and the test set. For each outer loop iteration, the outer loop training set can be re-split into a training set and a validation set using group-stratified cross-validation. Then, an "inner loop" of five training / validation iterations can be run based on the split outer loop training set to find a set of hyperparameters that optimizes performance. The entire outer loop training set can then be used to subsequently train a model using the optimized hyperparameters, which can be tested on an unknown outer loop test set. This process can be repeated once for each of the five different outer loop test sets, and the results of the optimized models in all outer loop test sets are summed to produce a result for the overall model performance. The variability of performance across the five outer loop test sets gives a measure of the generalizability of the model.
[0258] The nested cross-validation scheme for training and evaluation has two distinct advantages. First, it maximizes the amount of data tested, as each patient is accurately placed in the unknown outer loop test set once, which is particularly important for severity estimation tasks where there may be a small number of patients in the dataset. Second, nested cross-validation reduces the risk of bias in the optimized model due to tuning hyperparameters and testing the optimized model on the same test set.
[0259] As noted, it is proposed to provide severity estimates by training a machine learning classifier to distinguish between participants with different severities of respiratory disease (e.g., mild COPD or severe COPD) and using the probability of severe disease output by this classifier as an indication of COPD severity.
[0260] This "severity index" will indicate where the patient falls on the spectrum between mild and very severe COPD. When testing patients of all severity levels, it is expected that the output probability distribution for patients between mild and severe will fall between the distributions for mild and severe, and that the probability output (severity index) for a patient will steadily increase as their disease progresses.
[0261] In cases where there are a small number of patients available for this analysis, and especially in cases where there is a high prevalence of comorbidities among patients with mild disease, the analysis can be performed after removing patients with comorbidities from the dataset. This is because the machine learning model will be able to more effectively infer the underlying signal of COPD severity without the potential confounding of complications.
[0262] In this analysis, logistic regression can be used for classification due to its simplicity, interpretability, and because it has been shown to be as accurate as nonlinear models in the COPD classification task.
[0263] One objective was to apply interpretable machine learning techniques to capnography data and evaluate the performance of diagnostic and severity classifiers. Another objective was to construct a classifier that could distinguish capnography waveforms from patients with varying degrees of COPD severity from those without COPD. A second objective was to develop a severity index as an alternative to % predicted FEV1 that could be used by clinicians as an aid in quantifying COPD severity.
[0264] In summary, the techniques presented here facilitate real-time geometric waveform analysis and machine learning-based classification to diagnose the full range of severities of different respiratory diseases, such as COPD. In contrast to commonly used "black box" machine learning approaches, a set of highly interpretable methods is available that can provide a machine diagnosis back to individual geometric features of the pCO2 waveform and its associated physiological properties that suggest obstructive airway disease. Furthermore, the probabilistic output of the individual machine learning models can be used as a surrogate severity index for COPD, where the interpretability of the implemented machine learning techniques provides a picture of the capnographic progression of COPD patients from mild to severe.
[0265] Average waveform
[0266] Another approach to extracting features from one or more capnograms generated by a user is to use multiple respiratory waveforms to generate a single, smoothed average respiratory waveform, from which features can then be extracted. Extracting features from this average respiratory waveform can be performed instead of or in addition to extracting features from individual respiratory waveforms.
[0267] First, multiple respiratory waveforms are obtained from one or more capnograms generated by the user. While these respiratory waveforms can be provided as input to the method, in many cases, the method will include processing one or more capnograms to obtain the multiple respiratory waveforms. This typically involves dividing the one or more input capnograms into multiple capnogram segments. Preferably, the capnogram segments will each represent a single respiratory waveform corresponding to a single respiratory cycle, and the method will therefore include determining the delta and gamma turning points for each capnogram segment in any of the manners described above, and then extracting the portion of the capnogram segment between the delta and gamma turning points to generate the respiratory waveform. Alternatively, the cutoff point can be changed to capture more baseline, which can allow for the capture of more information related to the user's respiratory cycle. Alternatively, only a portion of the user-generated capnogram can be captured for the respiratory waveform, so that the respiratory waveform captures a portion of the respiratory cycle rather than the entire respiratory cycle.
[0268] Multiple respiratory waveforms will typically have different durations, reflecting the fact that a user's respiratory cycles will vary in length from breath to breath. Similarly, the CO2 of the breath will vary between respiratory cycles, resulting in different respiratory waveforms with different amplitudes. Because these differences between a single user's respiratory waveforms are generally not indicative of cardiopulmonary disease, it is preferred that the average respiratory waveform not reflect differences in the duration, and optionally, differences in the amplitude, of the multiple respiratory waveforms. Instead, in embodiments of the present invention, the average respiratory waveform is used to represent the average shape of a user's respiratory waveform.
[0269] To this end, the duration and, optionally, the amplitude of a plurality of respiratory waveforms are normalized to produce a plurality of normalized respiratory waveforms. The duration of a respiratory waveform refers to the time from the start of a respiratory cycle to the end of the respiratory cycle. As will be discussed below, the start and end points can be determined in different ways, but in all cases, these points are selected consistently across all respiratory waveforms. When normalizing, the duration of each of the plurality of respiratory waveforms is scaled so that the duration of all respiratory waveforms is the same.
[0270] As with the normalization of the durations of the multiple respiratory waveforms, the amplitudes of the multiple respiratory waveforms can optionally be scaled so that all respiratory waveforms have the same amplitude. This amplitude can be determined in various ways and can be as simple as the difference in CO2 readings between the highest and lowest points of the respiratory waveform. Ideally, the highest CO2 reading would be the end-tidal CO2 value, and this is indeed the case for many respiratory waveforms. However, due to measurement errors or underlying conditions of the corresponding respiratory cycle, the maximum amplitude is often found at different points in the respiratory waveform. This can lead to inconsistent ways of normalizing the respiratory waveforms, so it is preferable to use the end-tidal CO2 value to normalize the respiratory waveforms. This includes extracting the end-tidal CO2 value from each respiratory waveform using any of the methods described above (e.g., determining the beta turning point), and then adjusting the amplitude of each of the multiple respiratory waveforms so that each of the multiple respiratory waveforms has the same end-tidal CO2 value.
[0271] When the respiratory waveforms have been normalized, a generalized additive model (GAM) is preferably used to generate an average respiratory waveform from multiple normalized respiratory waveforms. While other averaging methods, such as taking the geometric mean or median of multiple respiratory waveforms, are possible, using a GAM allows for a more physiologically accurate waveform to be generated that does not overfit the data from multiple respiratory waveforms. It is also more resilient to outliers in the data.
[0272] Another advantage, which will be discussed in more detail with reference to the specific examples listed below, is that typically, each of the multiple respiratory waveforms comprises a series of data points. Simply taking the geometric mean or median of the data points from the multiple respiratory waveforms would result in an average respiratory waveform that itself comprises a discontinuous set of data points. In contrast, the use of a GAM results in a continuous functional representation of respiration, thereby allowing the average respiratory waveform to be used at any resolution desired by the operator.
[0273] However, in some embodiments of the present invention, other methods are used to generate the average waveform.
[0274] Features can then be extracted from the average respiratory waveform in a manner similar to that used to extract features from individual respiratory waveforms, and a trained machine learning model can be applied to these features to predict cardiopulmonary disease. A machine learning model can be used that utilizes features extracted from the user's individual respiratory waveforms as well as features extracted from the average respiratory waveform when predicting cardiopulmonary disease. Alternatively, a machine learning model can be used that utilizes features extracted from the user's individual respiratory waveforms or features extracted from the average respiratory waveform when predicting cardiopulmonary disease.
[0275] The usefulness of the generated average respiratory waveform, whether generated using GAM or some other method, can still be further improved by first identifying whether any of the multiple normalized respiratory waveforms is abnormal and then excluding the abnormal normalized respiratory waveforms before generating the average respiratory waveform.
[0276] One advantageous way to identify abnormal normalized respiratory waveforms is to interpolate the normalized respiratory waveforms so that they all have the same number of data points and directly compare the data points of one normalized respiratory waveform with the data points from another normalized respiratory waveform. Any normalized respiratory waveform with one or more abnormal data points is then identified as abnormal and excluded.
[0277] This approach has been found to be less computationally intensive than other methods of excluding abnormal breathing. For example, interpolating normalized respiratory waveforms and comparing data points to identify abnormal breathing has been found to be less computationally intensive than standard anomaly detection algorithms.
[0278] A specific example of this approach will now be described, which has been found to be particularly advantageous.
[0279] This example has been described in the context of obtaining a respiratory waveform from a single capnogram, but it can also be applied to the context of obtaining a respiratory waveform from multiple capnograms. Similarly, those skilled in the art will appreciate that other modifications to this specific example are possible. For example, the example described below can be modified to use respiratory waveforms collected from different portions of the respiratory cycle, rather than the waveforms between the delta turning point and the gamma turning point, such as those that collect more baselines or those that collect a predetermined portion of the respiratory cycle rather than the entire respiratory cycle.
[0280] First, a capnogram is obtained and then divided into M capnogram segments. Thereafter, for each capnogram segment, only the waveform between the δ turning point and the γ turning point is extracted. These turning points can be identified in any of the above-described ways, such as by extracting the waveform between the γδ angle starting point and the γ angle ending point of each capnogram segment. Using the δ and γ turning points is particularly advantageous because it results in an average respiratory waveform that is more similar to a single respiratory waveform extracted from the capnogram. In this way, features that can be extracted from a single respiratory waveform can also be extracted from the average respiratory waveform. Furthermore, in embodiments using a machine learning model, these points have been identified in one or more respiratory waveforms, which utilize features extracted from the average respiratory waveform in addition to the features extracted from the user's individual respiratory waveforms when predicting cardiopulmonary disease.
[0281] The respiratory waveforms are numbered from 1 to M, where respiratory waveform m is defined as the ,in For vector A vector of timestamps for each CO2 partial pressure value in , and with length All breaths are then normalized individually to have the same duration and, optionally, height (which is understood to mean the same amplitude), while preserving their unique shape. Optional height normalization can be achieved by scaling the end-tidal CO2 values (which can be calculated in any of the ways described above) to 5 kPa, so . Shift the time series points so that they start at 0, then linearly scale them so that the final point is at 3 seconds, . It will of course be understood by those skilled in the art that other values for the scaled end-tidal CO2 values may be chosen, as well as other values for the scaled duration. As mentioned above, although optional, in some cases it may be advantageous to exclude abnormal normalized respiratory waveforms. In this example, this is achieved by linearly interpolating all (normalized) respiratory waveforms to make them of the same length (that is, all respiratory waveforms include the same number of data points) to give and Those skilled in the art will appreciate that other values of N can be chosen. This allows the nth data point of each (normalized) respiratory waveform to be directly compared with the nth data point of every other (normalized) respiratory waveform. Respiratory waveforms in which one or more of these data points are found to be significant outliers can then be excluded as anomalous.
[0282] In this example, exclude those with at least one data point Any respiratory waveform , the data point The subset of normalized respiratory waveforms from the original group that is greater than 3 standard deviations from the median (i.e., the median of the nth data point of each respiratory waveform) is retained, i.e. Those skilled in the art will appreciate that different thresholds for abnormal respiratory waveforms may be selected. For example, a different number of standard deviations may be selected as the threshold, or a different metric may be selected, such as a percentage difference from the median.
[0283] Then use this group , then a generalized additive model (GAM) is used to generate a single, smoothed average breath of the capnogram. In this specific example, the GAM is defined as , which aims to minimize
[0284] The final term provides a cost function for the excess curvature, and the parameter The penalty level due to excessive curvature is thus controlled, thereby limiting the degree of overfitting. It will be appreciated by those skilled in the art that this last term is optional and can be set to Although other procedures such as Gaussian process regression or kernel ridge regression can be used, these procedures are more computationally intensive or produce less smooth curves and take longer to fit the high-resolution average respiratory waveform than GAM.
[0285] The function to be minimized can be obtained by Concatenate into a single collection (where the length of each vector is , and rewrite it by minimizing the following formula,
[0286] This model uses the breathing set Learning transformation functions Perform fitting.
[0287] To fit the model, first select A class of functions that should belong to this class. In this case, this is a set of cubic spline functions. These are a set of piecewise cubic polynomial functions that interpolate between specified control points (knots) while matching the zeroth, first, and second derivatives at those knots to form a single, smooth, continuous fit of the second derivative (as shown). Use a linear combination of a set of basis splines To build a cubic spline function, to give
[0288] The coefficient By minimizing and Perform fitting. The specific forms of are known in the art and are evenly distributed in time.
[0289] number is specified, where the first and last basis functions are positioned at the first and last times so that they can both be horizontally scaled to fit within the 0 to 3 second window. (As mentioned above, the respiratory waveform can be normalized to a different duration, in which case the basis functions can also be scaled to a different duration.) The coefficients are fitted using Poisson iteratively reweighted least squares (PIRLS) iterations. It has been found that This allows any type of breath (ie healthy or unhealthy) of any length to be fitted well. However, one skilled in the art will of course appreciate that different fitting and optimization procedures may be used.
[0290] The final digital average waveform output is ,in is A series of time points within a range, in this case meaning (As mentioned above, different values of N can be chosen.) Using this averaged respiratory waveform, features can be calculated using the methods described above.
[0291] The advantage of using spline functions is that this approach is computationally more efficient than other methods, such as fitting higher-order polynomials. Similarly, cubic basis splines have been found to represent the best balance between ensuring continuity of the GAM up to the second-order derivative and computational complexity. Lower-order splines are discontinuous at the second (and possibly first) derivative, and this curvature discontinuity reduces the effectiveness of feature extraction, as some features are based on curvature parameters of the mean respiratory waveform. Higher-order splines do not have this problem, but they increase the computational complexity of fitting the GAM and may also result in oscillations in the mean waveform that do not reflect the underlying data.
[0292] An alternative is to use Bezier curves instead of splines. However, using Bezier curves without knots is more computationally intensive, while using knots with Bezier curves conversely will result in discontinuities in the first-order derivatives. Both of these are significant disadvantages compared to the method of the present invention.
[0293] As already mentioned, other modifications can be made to the examples given above. One modification concerns how the breathing is normalized.
[0294] Although the duration and height of a respiratory waveform can be normalized to arbitrary values, in some preferred embodiments, respiration is normalized to the mean end-tidal CO2 value of a series of respiratory waveforms, where the mean end-tidal CO2 value is defined as follows:
[0295] This results in the following normalized height for each respiratory waveform:
[0296] Similarly, the duration of each respiratory waveform can be normalized to the average duration of the series of respiratory waveforms. This is particularly advantageous when the normalization step is combined with extracting the respiratory waveform from its corresponding capnogram segment, rather than receiving the respiratory waveform after the extraction occurs.
[0297] In this method, the timestamp of the delta angle is identified for each respiratory waveform. and the timestamp of the γ angle ,These The difference between defines the length of the portion of the respiratory waveform to be used when generating the average respiratory waveform. These values are then averaged to determine the mean delta time, mean gamma time, and mean difference:
[0298] (Although the arithmetic mean has been used as the average, other measures of the mean could be used, as could the median.)
[0299] Then, the timestamp of each respiratory waveform is converted as follows:
[0300] Here, the first term transforms and rescales the respiratory waveform so that the delta turning point is at 0 seconds and the time difference between the delta and gamma turning points is equal to the mean The second term then converts the breath back so that the delta turning point for that breath is at the timestamp of the mean delta turning point. (As mentioned above, the breath waveform can be normalized to any arbitrary value so that the length of 3s is ).
[0301] At this point, the delta turning point and the gamma turning point are consistent across respiratory waveforms. However, each respiratory waveform will typically include a portion of the baseline, and the amount of this baseline will vary between respiratory waveforms. This means that the start and end times of the respiratory waveforms will differ. To address this issue, the breaths are cut into predetermined lengths. This can be arbitrary, but to extract the breaths between the delta turning point and the gamma turning point, the following steps are preferably included.
[0302] First, identify the average duration of the breath from the gamma turning point to the end of the breath. For example, where the arithmetic mean is used to define the mean difference:
[0303] in The end of breathing timestamp.
[0304] Each respiratory waveform is then truncated to the nearest timestamp to define , where as mentioned above, (Alternatively, each respiratory waveform can be truncated to the mean respiratory duration, such that (when the arithmetic mean is used) . ) Any breaths that are smaller than the chosen intercept length are extended by copying their final value at regular intervals up to the intercept length, or by linear fitting up to the chosen intercept length.
[0305] Another modification to the example presented above involves how the optional step of eliminating abnormal breaths is performed. In some embodiments, the respiratory waveforms are interpolated to maintain the same length without eliminating abnormal breaths. In other embodiments, no interpolation is performed, and no abnormal breaths are eliminated. In those embodiments where the respiratory waveforms are interpolated and abnormal breaths are eliminated, it may be advantageous to restore the non-eliminated breaths to their original, non-interpolated form after eliminating the abnormal breaths.
[0306] The generation of the average waveform can also be done differently from the above methods. Although it is advantageous to use a GAM for this, this does not require the use of splines to learn the transformation function. For example, linear terms and factor terms can also be used. While it has been found that linear terms by themselves generally do not produce a good fit, factor terms lead to a better fit, although this fit is generally more jagged than when using splines. Thus, any combination of spline, factor, or linear terms can be used.
[0307] For example, when these additional terms are included, the function fitted by the GAM is
[0308] in of and of is the coefficient, and When fitting the transformation function The basis function of the factor term is an arbitrary value chosen when . and the basis functions of the linear terms Of course, it is known to those skilled in the art. and , that is, the weight assigned to the spline term is much larger than the weight of the linear term or the factor term, because the spline term provides a better fit.
[0309] It is also possible to generate an average waveform without using a GAM.
[0310] One such example uses the average of the CO2 values at each time point to produce an average waveform. For example, when using the arithmetic mean:
[0311] And, in addition, the standard deviation can be calculated at each time point:
[0312] As mentioned above, two other possible methods are Gaussian process regression and kernel ridge regression, although they are more computationally intensive or produce less smooth curves than GAM.
[0313] Similar to the GAM method, these methods begin by preprocessing the received breaths to obtain the respiratory waveform Then, kernel ridge regression aims to find the weight vector , the weight vector Minimize the following:
[0314] in is the loss function and Is to use the kernel function Constructed kernel matrix , where for example is the RBF kernel. Similar to the GAM implementation, Used as a smoothing parameter when the optimization algorithm tunes the parameters , you can set the smoothing parameter.
[0315] Then, use Generates an averaged waveform where yes A series of time points within a range.
[0316] Preprocess the received breath to obtain the respiratory waveform After that, the starting step of Gaussian process regression is to define the new time points of the average waveform as A series of time points within a range, which is referred to below as .
[0317] Gaussian process regression models will be known to those skilled in the art, and in this embodiment, a kernel function similar to that specified for the kernel ridge regression model is used, which will form a kernel matrix whose elements are defined as
[0318] Using Hyperparameters , the average waveform is predicted to be
[0319] and the standard deviation given by the vector is
[0320] (where square roots are element-wise) and is the identity matrix.
[0321] An additional optional but preferred step is to automatically select hyperparameters, which are those of the kernel function and This is achieved by minimizing the negative log marginal likelihood using a standard optimization algorithm:
[0322] The hyperparameters found from this minimization can then be plugged into the above calculation.
[0323] This variation of the Gaussian process regression framework is described as sparse Gaussian process regression and is computationally more efficient, although it sacrifices some degree of accuracy.
[0324] In obtaining After that, we then specify the "induced inputs". These induced inputs are smaller than Additional We will name these . We then define a set of kernel matrices using the elements of the induced input:
[0325] Further distribution parameters are also defined as follows:
[0326] These parameters allow
[0327] calculate.
[0328] (Note that the inductive input can be selected manually or, alternatively, automatically using an optimization algorithm , the optimization algorithm makes the following expression relative to Minimize:
[0329] in, , is normally distributed, and is the trace function.)
[0330] A final alternative to using GAMs is a procedure similar to Gaussian process regression, but which models each time point as a "Student's t" distribution rather than a Gaussian distribution.
[0331] In obtaining After that, the kernel function is now defined as before, but with the addition of a small variable , as shown below:
[0332] This gives the kernel matrix defined by the elements:
[0333] This gives the formula for calculating the average waveform and its standard deviation:
[0334] in are hyperparameters. As mentioned above, hyperparameters, including any hyperparameters from the kernel function, can first be manually selected or optimized by minimizing the negative log marginal likelihood as follows:
[0335] in ,and is the gamma function.
[0336] Any or all steps of the method can be performed in a remote or cloud computing device, or locally at the edge, i.e., on a device capable of retrieving capnography data. In one example, training is performed centrally prior to implantation or testing using a training model stored locally or on the device, for example, using model parameters stored in memory on an SoC or local computer. The trained model and its associated parameters can be stored centrally, i.e., in the cloud. The methods and processes described herein can be embodied as code (e.g., software code) and / or data. The models, methods, and algorithms can be implemented in hardware or software as is known in the field of machine learning. For example, hardware acceleration using a specially programmed graphics processing unit (GPU) or a specially designed field programmable gate array (FPGA) can provide certain efficiencies. For completeness, such code and data can be stored on one or more computer-readable media, which can include any device or medium capable of storing code and / or data for use by a computer system. When a computer system reads and executes the code and / or data stored on the computer-readable medium, the computer system executes the methods and processes implemented as embodied in the data structures and code stored within the computer-readable storage medium. In certain embodiments, one or more steps of the methods and processes described herein can be performed by a processor (eg, a processor of a computer system or a data storage system).
[0337] In general, any functionality described herein or shown in the accompanying drawings can be implemented using software, firmware (e.g., fixed logic circuitry), programmable or non-programmable hardware, or a combination of these implementations. As used herein, the term "component" or "function" generally refers to software, firmware, hardware, or a combination thereof. For example, in the context of a software implementation, the term "component" or "function" can refer to program code that performs a specified task when executed on one or more processing devices. The illustrated separation of components and functions into distinct units may reflect any actual or conceptual physical grouping and allocation of such software and / or hardware and tasks.
Claims
1. A method for classifying one or more capnograms generated by a user, the method comprising: obtaining a plurality of respiratory waveforms from one or more capnograms generated by the user; normalizing the durations of the plurality of respiratory waveforms to generate a plurality of normalized respiratory waveforms; generating an average respiratory waveform from the plurality of normalized respiratory waveforms; extracting features from the average respiratory waveform; and Applying a trained machine learning model to the extracted features, wherein the trained machine learning model is configured to output a classification of the one or more capnograms based on the features.
2. A method for training a machine learning model to learn indicators of a capnography waveform, the method comprising: For each average respiratory waveform, multiple average respiratory waveforms are generated by the following steps: obtaining a plurality of respiratory waveforms from one or more capnograms generated by the user; normalizing the durations of the plurality of respiratory waveforms to generate a plurality of normalized respiratory waveforms; and generating the average respiratory waveform from the plurality of normalized respiratory waveforms; extracting features from each of the plurality of average respiratory waveforms; obtaining a label for each average respiratory waveform indicating the capnogram classification of the corresponding one or more capnograms; as well as Using the extracted features and corresponding labels of the plurality of average respiratory waveforms, a machine learning model is trained to learn indicators of the capnography waveform to create a classification function. 3 . The method of claim 1 , wherein the average respiratory waveform is generated from the plurality of normalized respiratory waveforms using a generalized additive model (GAM). 4 . The method according to claim 1 , further comprising normalizing the amplitudes of the plurality of respiratory waveforms to generate the plurality of normalized respiratory waveforms.
5. The method of claim 4 , wherein normalizing the amplitudes of the plurality of respiratory waveforms comprises: Extract end-tidal CO2 values from each respiratory waveform; The amplitude of each of the plurality of respiratory waveforms is adjusted so that each of the plurality of respiratory waveforms has the same end-tidal CO2 value.
6. The method of claim 5 , wherein, for each of the plurality of respiratory waveforms, extracting the end-tidal CO2 value from each respiratory waveform comprises determining: The beta turning point between the expiratory plateau and the inspiratory descending limb.
7. The method according to any one of claims 1 to 6, wherein determining the beta turning point comprises: performing peak detection to identify local maxima of the respiratory waveform; Identify significant maxima from local maxima; When only a single significant maximum is identified, it is determined to be the beta turning point; and When multiple significant maxima are identified, the most significant maximum is determined and identified as the beta turning point.
8. The method according to any one of claims 1 to 7, wherein obtaining the plurality of respiratory waveforms comprises: dividing the one or more input capnograms into a plurality of capnogram portions, each capnogram portion representing a single respiratory waveform corresponding to a single respiratory cycle; as well as For each of the plurality of capnography portions: Determine the delta turning point between the expiratory baseline and the ascending limb of the exhalation; Determine the gamma turning point between the descending limb of inspiration and the inspiratory baseline; and A portion of the capnogram portion between the delta turning point and the gamma turning point is extracted to generate a respiratory waveform.
9. The method of claim 8, wherein determining the delta turning point comprises: determining a first time point at which a first-order differential of the respiratory waveform exceeds a delta threshold; and The first point is defined as the delta turning point.
10. The method according to claim 8 or claim 9, wherein determining the gamma turning point comprises: identifying a minimum of a first-order differential of the respiratory waveform; and, The gamma turning point is defined as the first time point after the minimum value where the first order differential of the respiratory waveform is above the gamma threshold.
11. The method of any one of the preceding claims, wherein determining one or more of the turning points comprises: The first order differential of the respiratory waveform is determined.
12. The method according to claim 11, wherein the method further comprises: comparing the first-order differential with a template; and, When the first-order differential of the respiratory waveform is inconsistent with the template, the respiratory waveform is excluded.
13. A method according to claim 11 or 12, wherein the first order differential of the respiratory waveform is determined using a Savitsky-Golay filter applied to the respiratory waveform.
14. The method of any one of claims 11 to 13, wherein determining one or more of the turning points comprises: A hump artifact is identified in the respiratory waveform and, when present, is processed during determination of the one or more turning points.
15. The method of claim 14, wherein identifying the hump artifact comprises: performing peak detection to identify local minima of the respiratory waveform; identifying a significant minimum from among the local minima; identifying a maximum value of the respiratory waveform and / or determining the beta turning point; dividing the respiratory waveform into a first portion and a second portion, the first portion and the second portion not overlapping, and the second portion including a maximum value of the respiratory waveform and / or the beta transition point; searching for a hump artifact in the first portion of the respiratory waveform when at least one significant minimum is identified; and / or When no significant minimum is identified, a first order differential of the respiratory waveform is used to search for a hump artifact in the first portion of the respiratory waveform.
16. The method of any one of claims 1 to 15, wherein generating an average respiratory waveform from the plurality of normalized respiratory waveforms comprises: identifying whether any one of the plurality of standardized respiratory waveforms is abnormal; and When generating the average respiratory waveform, the abnormal normalized respiratory waveform is excluded.
17. The method of claim 16, wherein identifying whether any one of the plurality of standardized respiratory waveforms is abnormal comprises: interpolating the normalized respiratory waveforms so that they all have the same number of data points; comparing data points of each interpolated respiratory waveform to data points from other interpolated respiratory waveforms; and Any interpolated respiratory waveform having one or more anomalous data points is identified as anomalous.
18. The method of claim 1, or any one of claims 3 to 17 when dependent on claim 1, wherein the classification of the capnogram includes a probability value corresponding to the severity of the respiratory disease.
19. A method for diagnosing cardiopulmonary disease, the method comprising the method of claim 1, or the method of any one of claims 3 to 18 when dependent on claim 1, further comprising: Diagnose the presence of cardiopulmonary disease.
20. The method of claim 2, or any one of claims 3 to 17 when dependent on claim 2, wherein the classification function is configured to predict the severity of a respiratory disease.
21. The method of claim 18 or 20, wherein the machine learning model is trained based on capnography waveforms annotated with corresponding severity of respiratory diseases.
22. The method of any preceding claim, further comprising obtaining a capnogram from the user, wherein the capnogram comprises a respiratory waveform.
23. The method of any of the preceding claims, wherein the machine learning model is trained and configured to predict the likelihood that the capnogram is associated with cardiopulmonary disease.
24. The method of any preceding claim, wherein the machine learning model is trained using labels belonging to a class indicating that capnography is associated with cardiopulmonary disease.
25. The method of claim 24, wherein the labels further belong to a plurality of categories, each category representing a capnography waveform associated with a cardiopulmonary disease, each category corresponding to a respective disease.
26. The method of claim 24 or 25, wherein the label further belongs to a category indicating that the capnogram waveform is not related to cardiopulmonary disease.
27. The method of any one of claims 23 to 26, wherein the machine learning model is configured to output a likelihood that the capnogram is associated with the cardiopulmonary disease.
28. The method of claim 27, further comprising stratifying the output by comparing the likelihood to a set of thresholds, each stratum representing a risk of the capnogram associated with the cardiopulmonary disease.
29. The method according to any one of claims 23 to 28, wherein the cardiopulmonary disease is selected from the group comprising COPD, asthma, asthma-COPD overlap syndrome (ACOS), small airways disease, the chronic bronchitis subtype of COPD, and the emphysema subtype of COPD.
30. Apparatus configured to carry out the method of any preceding claim.
31. A computer-readable medium comprising instructions which, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 to 29.
Citation Information
Patent Citations
capnometer
WO2017174983A1
Regulating device for a turbocharger
WO2019015800A1