Method and system for screening for obstructive sleep apnea during wakefulness using anthropometric information and tracheal breath sounds
Tracheal breath sound analysis with anthropometric data improves OSA screening accuracy to ~84%, addressing the limitations of existing methods by classifying OSA severity in subgroups, thus enhancing diagnostic efficiency and reducing costs.
Patent Information
- Application Number
- JP2022578733
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-18
- Filing Date
- 2021-06-17
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-06-17
AI Technical Summary
Current methods for diagnosing obstructive sleep apnea (OSA) are either costly and time-consuming (like polysomnography) or have low specificity (like the STOP-BANG questionnaire), necessitating a rapid, reliable, and objective screening tool with high sensitivity and specificity that can be applied during wakefulness.
A method using tracheal breath sound analysis during wakefulness, combined with anthropometric data, to classify patients as having or not having OSA by training a classifier with spectral and bispectral features, and applying a weighted voting scheme based on anthropometric subsets to improve accuracy.
Achieves an accuracy of ~84% with comparable sensitivity and specificity, effectively identifying OSA severity in subgroups with similar confounding variables, reducing the need for costly overnight assessments.
Smart Images

Figure 0007719107000015 
Figure 0007719107000016 
Figure 0007719107000017
Abstract
Description
[Technical Field]
[0001] The present invention relates to a computer-implemented system and method that utilizes machine learning to classify patients as having or not having obstructive sleep apnea based on recorded audio of the patient's breath, particularly breath sounds obtained during periods of full wakefulness. [Background technology]
[0002] Obstructive sleep apnea (OSA) is a common syndrome characterized by the recurrent occurrence of complete (apnea) or partial (hypopnea) pharyngeal collapse during sleep (American Academy of Sleep Medicine, 2005). The severity of OSA is generally measured by the apnea / hypopnea index (AHI), which is the number of apnea / hypopnea occurrences per hour of sleep. Usually, AHI < 5 is considered non-OSA, 5 < AHI < 15 is mild, 15 < AHI < 30 is moderate, and AHI > 30 is severe OSA. However, clinically, individuals with AHI < 15 are customarily considered people who may not benefit from treatment, and thus, AHI 15 is used as a threshold for determining severity (Littner, 2007; Medscape, August 23, 2016). The signs and symptoms of OSA include excessive daytime sleepiness, loud snoring, and the occurrence of breathing pauses, gasping, and / or choking during sleep (Epstein et al., 2009). OSA can significantly affect the quality of sleep and thus the quality of life. It is associated with an increased risk of cardiovascular disorders, hypertension, stroke, depression, diabetes, and headache, as well as an increased risk of traffic accidents (ResMed, 2013). If OSA is not treated, these complications may worsen (Bonsignore et al., 2019). In addition, it has been suggested that OSA consider the presence of serious complications in order to determine its severity and its treatment management (Vavougios, Natsios, Pastaka, Zarogiannis, & Gourgoulianis, 2016). Furthermore, failure to perform appropriate pre-checks before general anesthesia in OSA patients undergoing surgery can (due to the lack of accurate and reliable screening tools for OSA) result in perioperative morbidity and death. (American Society of Anesthesiologists, 2006; Gross et al., 2014) An accurate OSA screening tool will specifically reduce these risks before a patient undergoes surgery that requires general anesthesia.(American Society of Anesthesiologists, 2006; Gross et al., 2014). This paper reports on a new OSA classification procedure as a rapid and accurate screening tool based on anthropometric information and several minutes of respiratory sounds recorded during wakefulness.
[0003] The gold standard for diagnosing OSA is overnight polysomnography (PSG) assessment. However, it is expensive and time-consuming. Although many portable monitoring devices for OSA are available, they all require overnight recording. In Canada and the United States, approximately 10% of the population suffers from OSA (Young, T. et al., 2008), while the number of sleep rooms capable of performing PSG studies is limited. Therefore, there are long patient waiting lists, and in some locations, the wait time for a complete overnight PSG can exceed one year. Given these facts, having an efficient perioperative management plan for OSA based on an objective, reliable, and rapid diagnostic or screening tool is highly desirable, especially for anesthesiologists (American Society of Anesthesiologists, 2006; Chung & Elsaid, 2009; Gross et al., 2014).
[0004] A commonly used rapid OSA screening tool for patients undergoing surgery requiring general anesthesia is the STOP-BANG questionnaire (Nagappa et al., 2015). It is a simple, rapid, and inexpensive assessment method reported to have high sensitivity (~93%), but at the cost of very low specificity (~36%) (Nagappa et al., 2015). Any assessment method with low specificity indirectly increases the referral rate for full PSG investigations and, therefore, increases costs to the healthcare system. Therefore, there is a need for a rapid, reliable, objective technique with high sensitivity and specificity for OSA screening that can be applied during the awake state.
[0005] The use of tracheal breath sound analysis during wakefulness has also been proposed to screen for OSA (Elwali & Moussavi, 2017; Karimi, 2012; Montazeri et al., 2012; Moussavi, Z. et al., 2015). Several research groups worldwide have investigated the feasibility of using either tracheal breath sounds or vocal sounds during wakefulness to predict OSA (Goldshtein et al., 2011; Jung et al., 2004; Kriboy et al., 2014; Sola-Soler et al., 2014). Overall, these studies reported accuracies between 79.8% and 90%, with comparable sensitivity and specificity. While these accuracies are significantly superior to those of the STOP-BANG questionnaire, none of the reported accuracies represent the accuracy of blinded studies. In addition, their disproportionate sample sizes were very small relative to the diversity of the OSA population [23 (minimum 10 subjects for non-OSA) and 70 (minimum 13 subjects for OSA)].
[0006] Therefore, there is a need for improved methods and techniques for using analysis of fully awake tracheal breath sounds to screen for OSA patients. Summary of the Invention
[0007] According to a first aspect of the present invention, there is provided a method for providing an obstructive sleep apnea (OSA) screening tool, the method comprising: (a) obtaining an initial dataset comprising a respective subject dataset for each of a plurality of human test subjects whose respective audio recordings were taken during a period of fully awake state, the subject dataset comprising at least: the known OSA severity parameters of the study subjects; anthropometric data identifying different anthropometric parameters of the test subject; and audio data including stored audio signals from said audio recordings of each of said test subjects; and a process in which (b) extracting at least spectral and bispectral features from the speech data of the subject data set; (c) selecting a training dataset from the initial dataset, and from the training dataset, grouping together subject datasets representing a first high severity group of the test subjects, whose known OSA severity parameter is above a first threshold, and grouping together subject datasets representing a second low severity group of the test subjects, whose known OSA severity parameter is below a second threshold, the second threshold being less than the first threshold; (d) dividing the subject data set for each of the high severity group and the low severity group into a plurality of anthropometrically distinct subsets based on the anthropometric data; (e) at least in part, for each anthropometrically distinct subset, filtering the extracted features into a selected feature subset; To create a separate pool of predictive models for each anthropometrically distinct subset to predict OSA severity values. deriving inputs for a classifier training procedure by: (f) training a classifier using the derived inputs to classify human patients as either OSA or non-OSA based on a respective patient dataset, the patient dataset comprising at least: anthropometric data identifying different anthropometric parameters of the patient; and at least one of: audio data storing audio signals stored from audio recordings of each of said patients during a period of fully awake state; and / or feature data relating to features already extracted from said audio signals from said audio recordings of each of said patients. The method includes a step of storing the
[0008] According to a second aspect of the present invention, there is provided a method of screening a patient for obstructive sleep apnea (OSA), the method comprising: (a) obtaining one or more computer-readable storage media, the storage media including: a trained classifier derived in the manner detailed in steps (a) to (f) of the first aspect of the present invention; a patient dataset of the type detailed in step (f) of the first aspect of the invention for the patient; and storing statements and instructions executable by one or more computer processors; (b) through the execution of said statements and instructions by said one or more computer processors; (i) reading said anthropometric data of said patient data set; (ii) running said trained classifier multiple times, each time starting with inputs consisting of or derived from different combinations of suitable features consisting of at least spectral features and bispectral features selected or derived from said patient dataset, in particular for different anthropometric parameters read from said anthropometric data of said patient dataset, and deriving therefrom, for each run of said trained classifier, a respective classification result classifying said patient as either OSA or non-OSA; (iii) deriving a final classification result for the patient based on the classification result of step (b)(ii); and (iv) storing or displaying the final classification result; Includes.
[0009] According to a third aspect of the present invention, there is provided one or more non-transitory computer-readable media having stored thereon executable statements and instructions for performing step (b) of the second aspect of the present invention.
[0010] According to a fourth aspect of the present invention, there is provided a system for deriving or operating a screening tool for obstructive sleep apnea, the system comprising one or more computers, the one or more computers comprising one or more computer processors, one or more non-transitory computer-readable media connected to the one or more computer processors, and a microphone connected to one of the one or more computers for capturing audio recordings of each of a human test subject's and / or patient's respiratory cycle during a period of full wakefulness and storing the audio recordings as audio data on the one or more non-transitory computer-readable media, the one or more non-transitory computer-readable media also storing executable statements and instructions configured, when executed by the one or more computer processors, to extract features from the audio recordings and perform steps (a) to (f) of the first aspect of the present invention and / or step (b) of the second aspect of the present invention. [Brief explanation of the drawings]
[0011] Preferred embodiments of the present invention will now be described in conjunction with the accompanying drawings. [Figure 1] FIG. 1 is a schematic block diagram of a system of the present invention that can be used to build and subsequently execute a computerized screening tool for classifying patients as having or not having obstructive sleep apnea (OSA or non-OSA). [Figure 2] FIG. 2 illustrates schematically the recording procedure for recording tracheal breath sounds during periods of full wakefulness from both test subjects whose results will be used as training and blinded test data to construct a screening tool, and patients who will subsequently be screened using said screening tool. [Figure 3] 1 is a flow chart sequence illustrating the first and second data collection stages of subject data collection and the classifier training process for building the computerized screening tool described above. [Figure 4]10 is a flowchart sequence illustrating the third data processing and feature reduction / selection stage of the subject data collection and classifier training process. [Figure 5] 10 is a flowchart sequence illustrating a first portion of the fourth model creation and model selection stage of the subject data collection and classifier training process. [Figure 6] 10 is a flow chart sequence illustrating the remainder of the fourth stage of the subject data collection and classifier training process, followed by the fifth, model binding stage. [Figure 7] FIG. 7 is a flow chart sequence illustrating the first and second data collection stages of a patient selection process for screening patients using a computerized screening tool constructed from the data collection and classifier training processes of FIGS. 3-6. [Figure 8] 10 is a flow chart sequence illustrating the third data processing and patient classification stage of the patient selection process. [Figure 9] FIG. 9 shows a plot showing a typical respiratory inspiratory phase tracheal sound signal (lower graph) and its envelope (upper graph), which roughly represents its estimated flow rate, along with identification of the median time. [Figure 10] 7A-7C are schematic representations of simplified variations of the subject data collection and classifier training process of FIGS. 3-6. [Figure 11] Figure 11 shows the regression model of selected feature combinations for the age < 50 (top) and male (bottom) subsets from the first experiment using a simplified version of Figure 10, with log AHI, where the filled dots indicate which log AHI value was estimated by the model, AHI is the apnea-hypopnea index, and CC is the correlation coefficient. [Figure 12] From the first experiment, scatter plots from bag validation on the training dataset (top) and from the blinded test set (bottom) are shown, where circular and triangular plot points represent non-OSA and OSA individuals, respectively. [Figure 13]From the first experiment, the signals and average exponential spectra recorded from nasal inspiration are shown, in which the set of lines marked with triangles represents the OSA group, the set of unmarked lines represents the non-OSA group, and within both sets, the dotted lines represent the 95% confidence intervals. [Figure 14] 11 is a flow chart sequence illustrating the steps of the simplified subject data collection and classifier training process of FIG. 10. [Figure 15] Figures 3-6 show histograms of AHI responses before and after their scaling from a second experiment using the full subject data collection and classifier training process. [Figure 16] From the second experiment, scatter plots between responses and predicted responses are shown using three, four, and five feature models. [Figure 17] 3 is a schematic cross-sectional view of a microphone coupler having an air chamber (represented by hatching) particularly suited to the audio recording procedure of FIG. 2. [Figure 18] FIG. 18 shows the microphone chamber response during inspiration for various chamber dimensions of the microphone coupler of FIG. 17, with an enlarged inset showing the response curves from 50 to 500 Hz. [Figure 19] FIG. 18 shows microphone chamber response during exhalation for various chamber dimensions for the microphone coupler of FIG. 17, with an enlarged inset showing the response curves from 50 to 500 Hz. [Figure 20] Microphone chamber response for a chamber diameter of 10.00 mm during inspiration and expiration is shown, illustrating the effect of inclusion or omission of the body or vibration isolation rubber ring. DETAILED DESCRIPTION OF THE INVENTION
[0012] Systems and Processes Disclosed herein are details of a system and process that uses a novel, inventive decision-making algorithm, called AWakeOSA, that uses anthropometric information and several minutes of tracheal breath sounds recorded while awake to identify individuals with OSA who require treatment. The algorithm considers confounding anthropometric effects on breath sounds and uses them to predict OSA severity in subgroups with similar confounding variables.
[0013] Referring to Figure 1, tracheal breath sounds are preferably recorded using a flat-band, broadband microphone (10) inserted into a plastic microphone coupler that defines a cone-shaped air space between the microphone and the skin (Figure 1), e.g., approximately 2 mm deep. The microphone (10) is positioned above the suprasternal notch of the trachea and supported in this position using, e.g., double-sided adhesive ring tape. The acoustic signal is sent through a bandpass filter (12) and amplifier (14) and converted to a digital signal by an analog-to-digital converter (16), enabling digital recording of the sound by a computer (18) running appropriate recording software (20) and storing the digital sound data in computer-readable memory.
[0014] Referring to Figure 2, audio recordings were made separately with the subject / patient resting in both supine and upright positions, and with the subject / patient resting with their head both supported and unsupported by a pillow. The subject / patient was instructed to take multiple (>2) complete deep breathing cycles through their nose with their mouth closed, followed by multiple deep breaths through their mouth while wearing nose plugs. Recordings were either continuous or initiated with a silent recording period. All recordings began with the inspiration phase and were marked by the voice of the operator conducting the recording procedure or another audible marker clearly distinguishable from the recorded breath sounds. This same recording procedure was used for both "test subjects" whose recordings were used as training data for building a computerized OSA screening tool, and for "patients" whose recordings would be classified as either OSA or non-OSA using the trained screening tool.
[0015] Figure 3 illustrates the first and second stages of data collection and classifier training for building a computerized OSA screening tool. In step 101 of the first stage (100), the test subject's breathing sounds are recorded using the recording procedure described above in Figure 2. Next, in step 102, inhalation and exhalation sounds are extracted from the digitally recorded respiratory audio signals for each nasal and oral breathing maneuver and stored in separate files in computer-readable memory. Then, in step 103, all recorded signals are examined to exclude artificial signals, verbal noise, tones, interruptions, or respiratory phases or cycles with a low signal-to-noise ratio (SNR) compared to background noise. In step 104, any signals with fewer than two respiratory cycles are excluded from the analysis. Next, in step (105), using the logarithm of the variance of each phase signal (representing respiratory flow) (Yadollahi, A. & Moussavi, 2007), 50% durations surrounding the peripheral maximum of each respiratory phase are selected for further analysis, as shown schematically in Figure 9.
[0016] With further reference to FIG. 3 , in a second step (200), the information collected from each subject is added to a data store stored on the same or a different computer-readable memory as the digital audio recording. This stored information includes at least anthropometric data (201) indicative of OSA risk factors (e.g., BMI, age, etc.), and may also include responses to the STOP-BAND questionnaire (e.g., low / high blood pressure, snoring, etc.), and / or cranial and facial measurements (202) (e.g., facial width, neck width, neck depth, etc.). Various features and representations are calculated using the respiratory audio signal, such as fractal dimension and time domain, and are also recorded in the data store, as shown in (203). In step (204), each phase of the respiratory sound is individually filtered between 75 Hz and 3000 Hz to reduce the effects of heartbeat, muscle movement, distracting 60 Hz harmonics, and background noise.
[0017] Then, in steps 205-207, the power spectrum and bispectrum are calculated for each respiratory phase, and then the spectra are averaged for one maneuver (e.g., an expiratory phase recorded from the mouth) for each subject. In step 205, before calculating the power spectrum, each filtered signal is normalized by its variance envelope (i.e., a smoothed version of itself using dynamic averaging of a fixed sample sequence) (Gavriely & Cugell, 1995) and then by its energy (standard deviation) to remove the confounding effects of airflow fluctuations between respiratory cycles. In step 206, the power spectrum is calculated using Welch's method (Proakis, 2006) with a fixed window size and overlap between adjacent windows. The spectral and bispectral data are recorded in a data store in association with the stored information from the test subjects described above. The study subjects have already undergone or are being subjected to a PSG assessment, from which the AHI assigned to each subject is also stored in a data store, as shown at (208), in association with the anthropometric and craniofacial data (201), (202), and any other patient-specific information recorded for each subject, collectively referred to as the subject dataset.
[0018] The subject dataset recorded in the second stage (200) for the group of subjects forms an initial dataset, and data is drawn from this to perform the third stage (300) of the data collection and classifier training procedure. Referring to FIG. 4, first, in step (301) of this third stage (300), the subject dataset in the data store is divided into two severity groups: a non-OSA group with AHI < 15 and an OSA group with AHI ≥ 15. Next, as shown in (302A) and (302B), the entire initial dataset is divided into a training dataset and a blinded test dataset. Within the training dataset, for feature extraction, in (303A) and (303B), subject datasets with AHI ≤ 10 and AHI ≥ 20 are identified, where those with AHI ≤ 10 represent a low severity group and those with AHI ≥ 20 represent a high severity group. It will be understood that the specific AHI thresholds selected to identify the high and low severity groups of subjects may vary from these specific examples of AHI ≤ 10 and AHI ≥ 20. Records in the training dataset with 10 < AHI < 20 identified in (303C) are not used for feature extraction but are recombined with other members of the training dataset in step (307) for inclusion in the subsequent classifier training process described below.
[0019] As shown in (304A, 304B, 304C), the subject datasets are subdivided into anthropometrically distinct subsets (abbreviated "anthropometric subsets") within each of the training and blind datasets based on anthropometric information. Each anthropometric subset contains subject datasets within a certain anthropometric category of subjects (e.g., BMI<35, and age≦50, and >50, male, female, NC>40, MpS≦2, etc.). These anthropometric subsets preferably contain at least 20 individuals in each of the high and low severity groups. Within each anthropometric subset, and using the subject datasets for the high severity group with AHI≧20 and the low severity group with AHI≦10, various features are extracted from various signal representations (spectral, bispectral, and fractal dimension and time domain) in step (305). Preferred examples of extracted features are mean, standard deviation, entropy, distortion and skewness, centroid, Katz's fractal dimension, etc. For each anthropometric subset, features are extracted from the time and frequency domain analysis of the mouth and nasal breathing audio signals. The minimum bandwidth for feature selection may be set to 100 Hz.
[0020] In step (306), feature reduction is performed for each anthropometric subgroup based on, for example, feature significance, feature robustness, and feature correlation and redundancy. For feature significance, the p-value for each feature between the two OSA severity groups (AHI≧20 and AHI≦10) of the training dataset within the anthropometric subset is calculated using an unpaired t-test. Any features with a p-value >0.05 are excluded. For feature robustness, each feature is assigned a robustness score. Initially, all features have a robustness score of 0. Individuals from the two severity groups (AHI≧20 and AHI≦10) of the training dataset are used to create 15 OSA severity group subsets. These subsets are randomly generated and shuffled until all individuals have been selected at least once. All possible combinations of the generated subsets for each severity group (AHI≧20 and AHI≦10) of each training dataset are created. For each feature and each combination, p-values and normality checks are calculated separately for each subpopulation of the training dataset for the two severity groups (AHI ≥ 20 and AHI ≤ 10). The Lilliefors test (Lilliefors, 1967) is used for normality checks. If the p-value is ≤ 0.05 and the normality check is valid for each severity group (AHI ≥ 20 and AHI ≤ 10) in the training dataset, the robustness score of this feature is increased by one point. This process is repeated multiple times. Finally, each feature has an overall robustness score. Features with an overall robustness score of > 0.6 are selected for further analysis. Regarding feature correlation and redundancy, the available subject datasets for the two severity groups (AHI ≥ 20 and AHI ≤ 10) in the training dataset are used with a support vector machine classifier to calculate the training accuracy, specificity, and sensitivity for each feature in each anthropometric subset of the training dataset. All correlation coefficients between any two features are calculated. Any set of features with an intermediate correlation coefficient > 0.9 is removed except for the features with the highest training classification accuracy, specificity, and sensitivity. By the end of this stage, a final set of features is selected for each subset.
[0021] In step 307, the selected features for each anthropometric subset are evaluated using the full training data and the blinded test data set separately. Then, in step 308, the training data is used separately for each OSA severity group in each anthropometric subset to remove outliers (outside the adjacent values above and below the box plot) for each feature, and the adjacent values above and below the box plot are recorded (Benjamini, 1988; Tukey, 1977). In step 309, the recorded boundaries are used to remove any outside values from the blinded test features.
[0022] Figure 5 shows the fourth stage (400) of the data collection and classifier training procedure, in which the response / dependent variable / output should be a numerical or ordinal variable with the property that either an increase or decrease in its value increases the severity; therefore, the AHI is used as the response. As shown in step (401), the response should approximate a Gaussian distribution or be scaled to obtain a Gaussian distribution. First, the skewness of the response is checked; if its absolute value is greater than or equal to 1, the response is scaled; if the skewness is positive, its logarithm is taken; and if it is negative, its squared value is taken.
[0023] Next, in step (402), a linear model for the AHI is constructed using anthropometric measurements and acoustic features are constructed to preserve the interaction effects between the features used. Thus, a second-order interacting polynomial model of n variables (i.e., feature combinations) is generated for the AHI parameters; see Equation 1 for the polynomial model (X is the features used). A model is created for every feature combination. Each model generates a new array of values representing the AHI.
[0024]
number
[0025] Regarding the first step of model reduction in step (403), since the overall linearity may ignore some local non-linearities, the overall correlation between the prediction model and AHI may be spurious. Therefore, it is desirable to find the most linear model with less variance. Thus, in step (403), the range of the response variable is divided into multiple segments. These segments are divided based on the OSA severity category, for example, non-OSA (AHI < 5), mild OSA (5 < AHI < 15), moderate OSA (15 < AHI < 30), and severe OSA (AHI > 30). In this case, there are four segments, denoted as s1, s2, s3, and s4 respectively. The correlation coefficient is evaluated for each segment. Next, the four correlation values are averaged. Only models with a correlation coefficient exceeding a certain threshold, such as 70% of the maximum absolute mean correlation (i.e., the maximum of the absolute values of the mean correlations calculated for different models), are used for detailed analysis, that is, retained in the pool of selectable models that remain available for further selection and use in the subsequent steps of the process.
[0026] Next, in steps 404-409, segment overlap is assessed for the purpose of model reduction and selection. Steps 404A through 408A represent the first branch of this assessment, which is generated from an AHI perspective, i.e., using actual AHI values from each subject's subject data set, while steps 404B through 408B represent the second branch, which is generated from a model perspective, i.e., using predicted AHI values calculated for each subject by the model. Within each branch, the percentage of overlap is assessed between every two segments; in the four-segment example, this means there are a total of six assessments (i.e., s1-s2, s1-s3, s1-s4, s2-s3, s2-s4, and s3-s4). Note that s4 will have the most severely affected subjects, and s1 will have the least severely affected subjects. The overlap between segments can be viewed both from an AHI perspective (left branch in Figure 5) and a model perspective (right branch). Therefore, the evaluation is done from both perspectives in the preferred embodiment.
[0027] Looking at Figure 5 from the beginning to the left branch, in step (404A), for each model, the AHI predictions calculated by it are sorted based on their corresponding actual AHI values from the subject data set, and then divided into four segments. Next, in step (405A), the mean and standard deviation (std) of the model values are calculated for each segment separately. Until then, the boundaries of each segment are their mean ± std. In step (406A), m 4,p >m 3,p >m 2,p >m 1,p and all models without a mean (where m x,p is the mean value for segment x from viewpoint p) is rejected, i.e., removed from the pool. Using the estimated boundaries, the overlapping region for low severity segments is h -std h Although it is above, the overlapping area with respect to the high severity segment is mean L +std LThe subscripts H and L denote high and low severity segments, respectively. For the remaining models in the pool, if overlap exists, the percentage of subjects in the overlapped region relative to all subjects is calculated in step (407A). In step (408A), the average overlap percentage of the six assessments (overlap) is calculated.
[0028] Referring now to the right branch of Figure 5, in step (404B), for each model, the actual AHI values from the subject dataset are sorted based on the corresponding predicted values calculated by the model and then divided into four segments, with the number of subjects per segment equal to the number of subjects selected in step (404A) for each corresponding segment. Next, in step (405B), the mean and standard deviation (std) of the response values are calculated separately for each segment. Until then, the boundaries of each segment are their mean ± std. Steps (406B, 407B, and 408B) are performed in the same manner as described for the left branch of steps (406A, 407A, and 408A). In step (409), the average of the two average overlap percentages from the two perspectives of each model is calculated. In step 410, the remaining models in the pool are filtered to a subset of models whose average overlap is less than the other models, reducing the pool to, for example, the first 20 models with the lowest overlap percentage, which are selected and retained in the pool for further analysis.
[0029] In Figure 6, the fourth stage (400) continues at step (411), where the models selected in step (410) for each anthropometric subset are evaluated using the training data, which are then used to evaluate values representative of the blinded test dataset. In step (412), within each anthropometric subgroup, each selected model is validated using the training data and then tested using the blinded test data. For example, using an AHI of 15 as the threshold, the classification process is performed using a Random-Forest (RF) classifier (Breiman, 2001) with 1200 iterations, with each iteration using two-thirds of the training dataset for training and the other one-third for validation. The input for the classifier is the predicted AHI value of the selected model, thus representing a one-feature RF classifier. That is, for each iteration, the training dataset is randomly divided into a training set consisting of 2 / 3 of the training dataset and a test set consisting of the remaining 1 / 3 of the training dataset, and the following three evaluations are performed: 1) Training results a) training Input the predicted AHI values for 2 / 3 of the training data from the selected model, and b) Evaluation The predicted AHI values for the same 2 / 3 of the training data from the selected model are input and the output (OSA vs. non-OSA classification) is compared with the actual AHI values from the same 2 / 3 of the training data. 2) Verify the results a) training From the selected model, the predicted AHI values for 2 / 3 of the training data are input. b) Evaluation The predicted AHI values for the other third of the training data from the selected model are input and the output (OSA vs. non-OSA classification) is compared with the actual AHI values from the same third of the training data. 3) Test the results a) training From the selected model, the predicted AHI values for 2 / 3 of the training data are entered. b) Input the predicted AHI values for the blinded study dataset from the selected model and compare the output (OSA vs. non-OSA classification) with the actual AHI values from the blinded study dataset.
[0030] A cost matrix may be used to account for the difference in sample size between the two OSA severity subgroups (AHI ≥ 15 and AHI < 15) in the initial dataset. This procedure is repeated multiple times. The accuracy is then verified, and the blinded test results are averaged for each model to use as its evaluated performance metric.
[0031] In step (413), the pool of models is filtered or reduced by removing those with poorer performance evaluations, leaving a reduced number (e.g., 5) of models with the highest mean validation and blinded test accuracy, and the balanced validation sensitivity and specificity of these models are selected and recorded. For each of these models remaining in the pool, the mean validation and classification decision test for each subject is also recorded. At the end of the fourth stage (400), each model pool is compiled, evaluated, and reduced to the best group selected by their evaluated performance for the different anthropometric subsets, as appropriate.
[0032] Referring further to FIG. 6, the fifth stage (500) begins with step (501), which first tries different model combinations of the best selected models remaining in different model pools for different anthropometric subsets, where each model combination stores the results of each model from each anthropometric subgroup. Next, in step (502), a weighted classification decision per subject is calculated for each model, where the model's validation sensitivity is used as the weight if the participant is classified as severe (i.e., having OSA), and the model's -ve validation specificity is used as the weight if the participant is classified as normal (non-OSA). Next, the average classification decision between the different models for each model combination is calculated for each participating subject. Next, in step (503), the overall classification, sensitivity, and specificity are calculated for the validation and blinded testing results separately, and the best model combination that provides the highest accuracy for the validation and blinded testing with reasonable sensitivity and specificity is selected. This selected best model combination model is stored in computer readable memory along with the trained classifier module to form at least an essential part of a computer executable OSA screening tool that can be used to screen and classify patients as either OSA or non-OSA.
[0033] 7 and 8 illustrate a patient screening process for screening patients using a computerized screening tool, in which the tool's trained classifier module and best model combination are derived from execution of the subject data collection and classifier training process of FIGS. 3-6. Referring to FIG. 7, the first and second steps (1000, 2000) shown therein are substantially the same as the first and second steps (100, 200) shown in FIG. 3 for the subject data collection and classifier training process described above. The notable difference in FIG. 7 is that steps (1000, 2000) are performed for a patient whose AHI is unknown, as opposed to the known AHI of the test subject. Thus, the patient's actual AHI (208 in FIG. 2) is unknown in the context of the patient screening of FIG. 7 and is therefore not stored in a data store on a computer-readable memory. Otherwise, steps (1001)-(1005) and (2004)-(2007) of Figure 7, and stored data (2001)-(2003), are the same as steps 101-105 and 204-207 of Figure 3, and stored data 201-203, except that they are stored for "patients" rather than "subjects," and are therefore referred to herein as patient data sets.
[0034] In Figure 8, in the third stage (3000) of the patient screening process, anthropometric information from the patient dataset is first compared against the same anthropometric categories previously used to subdivide the training and blinded test datasets into anthropometric subsets in step 304 of Figure 3 above, thereby assigning each patient to a subset of anthropometric categories in step (3001). In step (3002), the best audio signal features for classifying the patient, determined as the basis for each best model, are then extracted and evaluated from the signal representations (spectrum (2006), bispectrum (2007), fractal dimension, and time domain (2003)) stored in the patient dataset. In this step, any outliers are removed according to the boundaries previously recorded in step 308. In step (3003), the extracted and evaluated features from the patient dataset are applied to a patient-matched model from the saved best model combination, thereby calculating an estimated AHI value for the patient. Next, in step (3004), for each patient-matched anthropometric model using the predicted AHI value, the classification decision (OSA vs. non-OSA) is evaluated using a trained random forest classifier, and then in step (3005), a weighted voting average of the classification results from the different patient-matched anthropometric models is calculated. The resulting weighted voting average represents the final classification decision for patient screening, which is recorded in the patient dataset in a data store and preferably displayed on a visual display on the computer 18 running the screening tool or another computer or device connected via a network.
[0035] The basic system architecture shown in Figure 1 will be described as being usable to perform both the subject data collection and classifier training procedures of Figures 3 through 6 and the patient screening procedures of Figures 7 and 8, although it will be understood that the computerized screening tools derived from the processes of Figures 1 through 6 may be stored on and executed on separate computers. Also, while a single computer is shown, it will be understood that executable software modules or other subcomponents are stored in computer-readable memory for execution by one or more computer processors to perform the various processes described herein, and that such software may be distributed across any number of computer-readable storage media for execution by any number of computer processors embodied in any number of computers interconnected by a suitable local or wide area network.
[0036] Experimental Support Experiment A: The use of tracheal breath sound analysis for screening for OSA was investigated. Breath sounds recorded from the mouth and nose were first sequestered into inhalation and exhalation, and then characteristic features were extracted by spectral, bispectral, and fractal analysis and subjected to a classification routine to estimate OSA severity.
[0037] It was hypothesized that upper airway (UA) deformation due to OSA would affect respiratory sounds even during a fully awake state, and that this effect should be detectable by tracheal breath sound analysis (Finkelstein et al., 2014; Lan et al., 2006; Moussavi, Z. et al., 2015). Previous studies (Elwali & Moussavi, 2017; Karimi, 2012; Montazeri et al., 2012; Moussavi, Z. et al., 2015) demonstrated this hypothesis. Subsequently, a test classification accuracy of ~84% was achieved with comparable specificity and sensitivity (<10% difference) for two balanced groups: non-OSA (AHI ≤ 5, n = 61) and OSA (AHI ≥ 10, n = 69) (Elwali & Moussavi, 2017). The use of tracheal breath sound characteristics to screen for OSA during a fully awake state has also been shown to have significant advantages over solely using anthropometric information (i.e., gender, age, and neck circumference) (Elwali & Moussavi, 2017). However, the effects of anthropometric confounding variables (i.e., age, gender, height, etc.) on the acoustic signal were not investigated in such prior studies. Additionally, because it is desirable to achieve rapid screening with high sensitivity to identify OSA individuals who require treatment, it is desirable to have a single threshold (e.g., AHI = 15) for such decision-making. Although it is difficult to distinguish between AHIs of 14 and 16, AHI = 15 was the most common clinically accepted threshold for separating OSA individuals who require treatment from those who would not benefit from treatment (Littner, 2007; Medscape, Aug 23, 2016). When this threshold of AHI = 15 was applied in prior art studies, blinded test accuracy decreased to an undesirable <70%. It is generally argued that the AHI is not the best indicator for determining the diagnosis of OSA. Sleep medicine physicians typically base their decision on several factors, including daytime sleepiness and the number of nighttime awakenings, along with the AHI.However, for rapid screening involving automated real-time screening, control is necessary for accuracy, and AHI is the most commonly used standard.Therefore, the solution of the present disclosure uses anthropometric information and several minutes of tracheal breath sounds recorded during a fully awake state to identify OSA individuals who need treatment.The AWakeOSA algorithm takes into account the confounding anthropometric effects that affect breath sounds and uses them to predict OSA severity in subgroups with similar confounding variables.
[0038] The premise of the AWakeOSA algorithm is to find the acoustic features with the best sensitivity to OSA severity (determined by AHI) for each subgroup of individuals with specific anthropometric factors (i.e., age, sex, weight, etc.). A simplified version of the algorithm tested in Experiment A is represented schematically in FIG. 10 and shown in more detail in the flowchart sequence of FIG. 14, which includes a feature reduction and selection step (300) similar to that of FIG. 4, but omits the subsequent model building, model selection, and model combination steps, the embodiments of which are shown in more detail in FIGS. 5 and 6. Because the OSA population is highly heterogeneous and many confounding variables, such as age, sex, height, weight, etc., affect the characteristics of respiratory sounds (Barsties, Verfaillie, Roy, & Maryn, 2013; Linville, 2001; Titze, 2000; Torre III & Barlow, 2009), it is difficult to predict AHI for all individuals using a few acoustic features. The simplified AWakeOSA algorithm overcomes this difficulty by grouping individuals into subgroups based on specific anthropometric factors that each affect respiratory sounds. Then, in each subgroup, the best acoustic features for predicting AHI are extracted, and a classifier is trained using the training dataset. The classifier results for each subgroup are then used in a weighted average voting scheme to make a classification decision: OSA (AHI > 15) or non-OSA (AHI < 15).
[0039] The following summarizes the procedures and results of a first experimental use of the simplified AWakeOSA algorithm, where the threshold AHI=15 and data were collected from 199 individuals with various OSA severities (AHI ranged from 0 to 143), with 45% of the individual data set aside for blinded testing and the remainder used to extract features and train a classifier. In all examples, a two-group classification algorithm based on the Random-Forest algorithm (Breiman, 2001) is used, although it will be understood that different types of classification algorithms may alternatively be employed in other useful embodiments of the present invention.
[0040] Respiratory sound data from 199 participants were used, consisting of 109 non-OSA individuals (50 men, AHI < 15) and 90 OSA individuals (66 men, AHI > 15). All data were recorded during full wakefulness in the evening (~ 8 PM) before participants proceeded to the PSG sleep study. While the detailed example specifically uses AHI and respiratory sounds, other embodiments may use OSA severity parameters other than AHI and / or alternative audio recordings such as speech. Anthropometric parameters for the two groups are reported in Table 1. Data from the two severity groups (AHI > 15 and AHI < 15) in the initial dataset were not matched on any of the following confounding variables: sex (male / female), body mass index (BMI threshold = 35), neck circumference (NC threshold = 40), age (threshold = 50), or Mallampati score (MpS threshold = 3). Of the 199-individual dataset, data from 86 individuals (47 non-OSA and 39 OSA) were reserved for blinded testing to evaluate the accuracy of the algorithm. Table 1 also shows the anthropometric information for the training dataset and the two severity groups (AHI ≥ 15 and AHI < 15) in the test dataset.
[0041] [Table 1]
[0042] AHI was found to have significant correlations with BMI, NC, and MpS (r = 0.44, 0.43, and 0.26, respectively). However, when available anthropometric information (i.e., BMI, age, NC, and sex) from the STOP-BANG questionnaire (Nagappa et al., 2015) was used and the entire dataset (199 subjects) of the two OAS severity groups (threshold AHI = 15) was classified, the resulting classification accuracy, specificity, and sensitivity were found to be only 63.4%, 74.3%, and 50.5%, respectively.
[0043] Recorded breath sounds collected from four breathing methods (i.e., mouth and nose inspiration, and expiration) were analyzed. A sharp threshold AHI = 15 was used to train the classifier, while for the feature extraction and reduction phase, only data from subjects with AHI ≤ 10 (n = 60) and AHI ≥ 20 (n = 40) in the training data phase were used. Using the algorithm outlined in Figure 10, recorded breath sounds were analyzed separately for each of the anthropometric subsets of the training dataset: BMI < 35, Age > 50, Age ≤ 50, male, NC > 40, and MpS ≤ 2. Subjects in each anthropometric subset were matched for only one anthropometric variable. Due to the limited sample size, subsets such as BMI > 35 or NC < 40 did not exist. The subset selected and used included 30 non-OSA subjects and 20 OSA subjects, which was a reasonable number for feature extraction, feature reduction, and group classification.
[0044] The number of respiratory sound features extracted from the two breathing techniques was approximately 250, while the inspiratory and expiratory phases were analyzed separately. Using a novel feature reduction procedure (Figures 4 and 14) on the training dataset, approximately 15 features per anthropometric subset were selected for further investigation. These sound features indicated significant differences between the two OSA severity groups (AHI > 15 and AHI < 15), as they were highly correlated with AHI (p < 0.01). Furthermore, the sound features showed an effect size > 0.8. Table 2 shows the definitions of the selected sound features, breathing techniques, investigated subsets, and their correlation coefficients with AHI.
[0045] [Table 2-1]
[0046] [Table 2-2]
[0047] In the next stage of the simplified AWakeOSA algorithm, the selected acoustic features and one anthropometric feature, NC, were used as features for classification because NC showed a significant correlation with AHI (0.43, p<0.01) when tested on the training dataset. The NC feature was further investigated in all subsets except for its own subset. Thus, the selected acoustic features and NC were divided into three features and four feature combinations. These feature combinations were used to classify each participant's data in all subsets into one of two OSA severity classes (AHI ≥ 15 and AHI < 15). Column 4 of Table 2 shows the acoustic features used for each anthropometric subset to construct the classification combination. NC was selected and used in the subsets of age ≤ 50 and males.
[0048] Table 3 shows the classification accuracy, specificity, and sensitivity of the out-of-bag validation using Random-Forest classification (Breiman, 2001) and blinded test datasets for each anthropometric subset, with the final classifications separately shown. Additionally, Table 3 shows the correlation coefficient between AHI / log(AHI) and each feature combination using linear regression analysis. Figure 11 shows the results of a linear regression analysis using AHI on a logarithmic scale for the feature combinations selected for the age ≤ 50 subset and the male subset. Furthermore, the feature combination selected for age ≤ 50 showed the highest test classification accuracy of 86% within its own subgroup. When this feature combination was used for the entire training dataset as well as the blinded dataset, it resulted in an out-of-the-bag validation accuracy of 71.6% for the training dataset and an accuracy of 75.6% for the blinded test data.
[0049] [Table 3]
[0050] The overall classification results using the proposed weighted linear voting scheme, illustrated in Figure 10, revealed classification accuracy, specificity, and sensitivity of 82.3%, 82.3%, and 81.4% for out-of-bag validation, and 81.4%, 80.9%, and 82.1% for blinded validation data, respectively. Figure 12 shows scatter plots of overall classification decisions for out-of-bag validation (top) and blinded validation (bottom). A classification decision of 1 (specificity and sensitivity both considered 100%) means that all used Random Forest classifiers for each anthropometric subset voted the subject into the OSA group, while a classification decision of -1 (specificity and sensitivity both considered 100%) means that all used Random Forest classifiers for each anthropometric subset voted the subject into the non-OSA group.
[0051] The anthropometric information of the misclassified subjects in both the out-of-bag validation and blinded test classifications is listed in Table 4. Of the 109 non-OSA subjects, 20 were misclassified in the OSA group, and of the 90 OSA subjects, 16 were misclassified in the non-OSA group. Investigations were also conducted to determine whether removing a subset of decisions from the voting stage in the simplified AWakeOSA algorithm improved or worsened the overall classification results. Results showed that removing any of the subsets reduced classification performance by 2.5–8%.
[0052] [Table 4]
[0053] Undiagnosed severe OSA significantly increases healthcare costs and the risk of perioperative morbidity and mortality (American Society of Anesthesiologists, 2006; Gross et al., 2014). Anthropometric measures have been shown to have high sensitivity in screening for OSA (El-Sayed, 2012; Nagappa et al., 2015), but at the expense of very poor specificity (~36%) (Nagappa et al., 2015). This may be due, in part, to the subjectivity of most STOP-BANG parameters. Although these parameters do not have good classification power for screening for OSA, they correlate with AHI and also affect respiratory sounds (Barsties et al., 2013; Linville, 2001; Nagappa et al., 2015; Titze, 2000; Torre III & Barlow, 2009). In the dataset from Experiment A, anthropometric parameters correlated with AHI (0.03 < |r| ≤ 0.44) and with breath sounds (0.00 < |r| ≤ 0.5). Breath sounds have previously been shown to have much higher classification power for screening OSA than anthropometric features (Elwali & Moussavi, 2017). When sound features alone were used to classify OSA severity with a sharp AHI = 15, the classification accuracies for out-of-bag validation and blinded testing were found to be 79.3% and 74.4%, respectively. These accuracies are much higher than those provided by the STOP-BANG questionnaire (63.4%), but are not sufficiently high for reliable (stable) OSA screening during a fully awake state.
[0054] To increase the accuracy and reliability of using breath sounds to screen for OSA during wakefulness, a simplified variant, shown diagrammatically in FIG. 10, subgroups the data and uses anthropometric features to generate classification votes in each subgroup, while the final classification is based on the average vote of all subsets. This subdivision was performed to reduce the influence of confounding variables on the feature extraction stage and the classification process. Results showed that the simplified AWakeOSA voting algorithm provided higher and more reliable blinded test accuracy, even at a sharp threshold AHI = 15. A major challenge in all studies using breath sound analysis to screen for OSA during wakefulness is population heterogeneity. Because the causes of OSA can vary between individuals, people with the same OSA severity may have very different anthropometric features, and these differences affect breath sounds differently. The simplified algorithm of FIG. 10 takes such heterogeneity into account and attempts to find the best sound features specific to each subset of data that share one particularly important confounding variable, such as age, BMI, or gender. Meanwhile, the variant of FIG. 10 allows for weighting factors for each subset classifier's vote, as these features affect sound variation. The weighting factors for each subset classifier's vote are based on the sensitivity and specificity of the classifier developed during the training phase.
[0055] Anthropometric subsets were divided based on age, gender, BMI, Mallampati score, and neck circumference. Age-related effects include a decrease in female voice pitch and an increase in male voice pitch (Torre III & Barlow, 2009). Furthermore, aging causes a decrease in muscle mass, dry mucous membranes, and an increase in speech variability (Linville, 2001). Thus, two anthropometric subsets were used for ages >50 and <50. Gender was selected as an anthropometric subset because female voices are known to have higher fundamental frequencies than male voices, which is unrelated to OSA severity (Titze, 2000). Thus, it is important to analyze male and female breath sounds separately. BMI significantly influences voice quality and breath sounds (Barsties et al., 2013). Thus, two anthropometric subsets were used for BMI >35 and BMI <35. The Mallampati score is an index of pharyngeal size and is therefore associated with upper airway narrowing (Gupta, Sharma, & Jain, 2005). Thus, MpS > 2 and MpS < 3 were selected to form two anthropometric subsets. Neck circumference is one of the most important OSA risk factors (Ahbab et al., 2013). Therefore, participants with NC < 40 and participants with NC > 40 were investigated separately. Given the larger dataset size, it may have been beneficial to form anthropometric subsets based on height and weight, which are independent of BMI, to allow for a larger number of subjects in each subset. However, BMI and NC are highly correlated with height and weight, and thus, weight- and height-based anthropometric subsets may not be necessary to significantly improve accuracy.
[0056] Among the anthropometric subsets, the younger age subset (age ≤ 50) showed a higher test classification accuracy (86%) compared to the others. This was not a surprising result. Age is a well-known risk factor for OSA (Bixler, Vgontzas, Ten Have, Tyson, & Kales, 1998). Healthy individuals aged ≤ 50 generally do not suffer from loss of upper airway muscle mass, but these individuals, who subsequently experience upper airway collapse during apnea / hypopnea events, have more muscle tone than OSA individuals of the same age. Thus, loss of muscle tone correlates with AHI and significantly affects breath sounds in this group. Furthermore, after examining the anthropometric information of the younger age subset, it was found that the majority of non-OSA individuals in this subset also had a BMI < 35, MpS < 3, NC < 40, and were female. Meanwhile, the majority of OSA individuals in this subset were from the opposite category. Therefore, the young age subset was expected to have the highest classification accuracy among the other subsets for classifying individuals into AHI>15 and AHI<15.
[0057] Another interesting observation about the combination of features selected for the younger and male subsets is their high correlation with AHI on a logarithmic scale (0.66 and 0.6, respectively), as shown in Figure 11. This suggests that OSA severity increases exponentially with increasing AHI. Clinically, this also presents a challenge in diagnosing OSA in individuals with relatively low AHI. Additionally, very high AHI coincides with overt clinical symptoms.
[0058] Among the selected sound features, only six features were extracted from frequency components above 1100 Hz, while the remaining features were extracted from frequency components below 600 Hz. This indicates that OSA most strongly affects low-frequency sound components. Thus, it is important to record respiratory sounds using equipment that does not filter either low or high frequencies. One of the best sound features was the mean slope of the power spectrum of the nose inspiratory signal recorded within the 250–350 Hz frequency band. This feature was selected three times in the subset of subjects with age ≤50, males, and MpS <3. During nasal inspiration, the muscles of the upper airway are active while the pharyngeal cavity is patent. The pharyngeal cavity is responsible for transporting air from the nose to the lungs and vice versa. The upper airway of individuals with OSA is generally characterized by pharyngeal narrowing, a thick tongue, loss of muscle tone, and a thick, elongated soft palate (Lan et al., 2006). These characteristics contribute to a significant narrowing of the upper airway lumen and increase the likelihood of developing OSA during sleep. The average slope of the spectrum from 250 to 350 Hz indicates that sound power increases after 250 Hz in OSA individuals, but decreases in non-OSA individuals (Figure 13). It also indicates that OSA individuals tend to have higher resonant frequencies than non-OSA individuals. These results indicate that the OSA group is characterized by a more distorted and stiffer upper airway than non-OSA individuals. This is consistent with MRI / CT imaging studies that have shown that the upper airways of OSA individuals during a fully awake state have, on average, greater regional compliance and stiffness (Finkelstein et al., 2014; Lan et al., 2006).
[0059] Based on the final overall classification decision (see Figure 12), classifying a subject with an overall classification decision >0.7 or <-0.7 has approximately 90% confidence in being classified into the correct class. The significant contribution of all subsets to ensuring the final vote for grouping was also investigated. Excluding one of the subsets selected from the final voting stage reduced the overall classifier performance. Therefore, the proposed subset was considered important for the analysis in this experimentally supported embodiment. It was interesting to note the anthropometric parameters of the incorrectly classified subjects. As can be seen from Table 4, the majority of incorrectly classified subjects in the non-OSA group were age <50, male, BMI >35, NC >40, and MpS <3. Meanwhile, the majority of incorrectly classified subjects in the OSA group were age >50, male, BMI <35, NC >40, and MpS <3. Bold indicates risk factors for the opposite group based on the STOP-BANG questionnaire. This suggests that although anthropometric parameters or risk factors are correlated with AHI, they do not have classification power and may result in misclassification.
[0060] The results of Experiment A above demonstrate that anthropometric parameters affect organ respiratory sounds and that this effect can be proactively utilized to improve the accuracy and reliability of OSA identification during awake states. All sound features selected for each anthropometric subgroup showed statistically significant differences between the two OSA groups (AHI > 15 and AHI < 15). Despite using a sharp AHI threshold (AHI = 15), the tested AWakeOSA algorithm demonstrated promising high classification-blind test accuracy with a sensitivity of 82.1% and a specificity of 80.9%. Thus, the AWakeOSA technology is expected to be of interest for OSA identification during awake states, especially for anesthesiologists preparing patients for general anesthesia. Such reliable and rapid OSA identification would significantly reduce perioperative resources and costs. It also helps reduce the number of undiagnosed OSA cases and the need for PSG evaluation, thereby significantly reducing medical costs.
[0061] The simplified algorithm employed in Experiment A incorporates all features from FIGS. 1 through 4 except for skull feature 202 (FIG. 3), but jumps directly from steps 308 and 309 to step 412 (random forest classification) without intervening steps 401-411 and 413-501 of the more comprehensive embodiment encompassing all of FIGS. 4 through 6. This more comprehensive embodiment is the subject of Experiment B and is described in more detail below. After step 412, the data collection and classifier training process of the simplified embodiment of Experiment A jumps to a weighted average voting step for the final classification decision, as in 502 of FIG. 6, and lacks the subsequent model combining step 503, since the model creation step and model selection steps 401-411 are omitted in the simplified variant.
[0062] In Experiment A, study participants (test subjects) were randomly recruited from those referred for overnight PSG evaluation at Misericordia Health Centre (Winnipeg, Canada). Recordings were conducted approximately 1–2 h before the PSG study. While fully awake and in a supine position (head resting on a pillow), tracheal respiratory sound signals were recorded using a Sony microphone (ECM77B), embedded in a small chamber placed over the suprasternal fossa of the organ, allowing for a ~2 mm space between the skin and the microphone. Participants were instructed to take five deep breaths for each maneuver, first through the nose at equal flow rates, and then through the mouth. At the end of each maneuver, participants were instructed to hold their breath for a few seconds to record background noise (called a silent period). Details regarding the recording protocol are available on the technology website (Elwali & Moussavi, 2017). AHI values were extracted from PSG recordings at Misericordia Health Center after overnight PSG evaluation prepared by a sleep specialist.
[0063] Of the 300 recorded respiratory sounds, data from 199 individuals were selected as valid data. The criteria for inclusion of an individual's data were that it had at least two clean (artifact-free, tone-free, interruption-free, and low SNR) respiratory cycles for each breathing technique. Each individual's sound signal was examined in the time-frequency domain by auditory and visual means to separate the inspiratory and expiratory phases. During our fully awake recordings, the first inspiratory phase was always observed to ensure 100% accurate separation of the two phases. This initial data set included data from 109 individuals with an AHI <15 and 90 individuals with an AHI >15. For simplicity, these two groups are hereafter referred to as the non-OSA and OSA groups. A block diagram illustrating the simplified methodology is presented in Figure 14. Anthropometric information for the analyzed data (199 individuals) is presented in Table 1.
[0064] As a whole, the data of 113 individuals were used for training, and the data of the remaining 86 individuals were used as a blinded test dataset. Within the training set, the data of subjects with AHI ≤ 10 (n = 60) and AHI ≥ 20 (n = 40) were used for feature extraction of two groups of non-OSA and OSA (AHI < 15 and AHI ≥ 15) respectively, and these data were from 100 subjects. The data of the remaining 13 subjects in the training dataset with 10 < AHI < 20 were not used for feature extraction but were included in the training classification process. Within the dataset of 100 subjects used for training, a body measurement subset was created, namely BMI < 35, age ≤ 50, and age > 50, male, NC > 40, and MpS ≤ 2. These subsets each had at least 30 individuals and 20 individuals in the non-OSA group and the OSA group respectively.
[0065] Spectral and bispectral features were extracted from the respiratory sound data. The signal power spectrum was calculated using the Welch method (Proakis, 2006), and the bispectrum was calculated using an indirect classification of the traditional bispectral estimator (Nikias & Raghuveer, 1987). Previously (Elwali & Moussavi, 2017), it was found that the frequency range could be divided into four main discriminant frequency bands (i.e., 100–300 Hz, 350–600 Hz, 1000–1700 Hz, and 2100–2400 Hz). Using these four frequency bands, spectral and bispectral features were extracted. Some features (i.e., mean, standard deviation, spectral entropy, skewness and skewness, spectral centroid, etc.) were extracted from the non-overlapping area between the mean spectrum / bispectrum and the 95% confidence intervals of the two severity groups of the training dataset (AHI ≥ 20 and AHI ≤ 10). The minimum bandwidth for feature selection was set at 100 Hz. As an example, Figure 13 shows the power spectra of inspiratory nasal breathing for the two severity groups (AHI ≥ 20 and AHI ≤ 10). The Katz and Higuchi fractal dimension (Higuchi, 1988; Katz, 1988) and the Hurst exponent (Korida, 1951) were also calculated from the signals in the time domain. Therefore, for each subset, approximately 250 features were extracted from the time- and frequency-domain analyses of the oral and nasal respiratory sound signals. All features were scaled to the range [0, 1].
[0066] The p-value of each feature was calculated between the two OSA severity groups (AHI ≥ 20 and AHI ≤ 10) in the training dataset using an unpaired t-test. Any features with a p-value > 0.05 were excluded. A robust score was assigned to each feature. Initially, all features had a robust score of 0. Subgroups of 15 individuals were created for each OSA severity group using the available individuals in the two severity groups (AHI ≥ 20 and AHI ≤ 10). These groups were randomly generated and shuffled until all individuals were selected at least once. All possible combinations of the subgroups generated for each severity group were created. For each feature, the p-value and normality check for each subgroup in the two severity groups (AHI ≥ 20 and AHI ≤ 10) using each combination were calculated separately. The Lilliefors test was used for normality checks (Lilliefors, 1967). If the p-value was ≤0.05, and the normality check for each severity group (AHI ≥ 20 and AHI ≤ 10) was valid, the robust score of this feature was increased by 1 point. This process was repeated 20 times. Ultimately, each feature had an overall robust score. Features with an overall robust score > 0.6 of the maximum robust score were selected for further analysis.
[0067] Using available individual data from two severity groups (AHI ≥ 20 and AHI ≤ 10) and a support machine classifier, training accuracy, specificity, and sensitivity were calculated for each feature in each anthropometric subset of the training dataset. All correlation coefficients between any two features were calculated. Any set of features with an inter-feature correlation coefficient ≥ 0.9 was removed, except for the feature with the highest training classification accuracy, specificity, and sensitivity. A final set of features for each subset was selected by the end of this phase. The efficacy of each selected feature was checked using Glass's delta equation (Glass, Smith, & McGaw, 1981). The features selected for each anthropometric subset were evaluated using the full training data of 113 individuals (62 with AHI < 15 and 54 with AHI > 15) and a blinded test dataset of 86 individuals (47 with AHI < 15 and 39 with AHI > 15). Using the training data for each OSA severity group for each anthropometric subset, outliers for each feature were removed, and the upper and lower flanking values of the boxplots were recorded (Benjamini, 1988; Tukey, 1977). Using the recorded values, the minimum lower flanking value and the maximum upper flanking value were recorded. Using the blinded test data, any values outside the recorded boundaries were removed.
[0068] Within each anthropometric subset, selected features were combined to create three-feature and four-feature combinations. Using each combination and the training data, a Random Forest (Breiman, 2001) classifier with 2 / 3 of the data in the training bag and 1 / 3 of the data removed from the bag was used to evaluate accuracy, specificity, and sensitivity in an out-of-bag validation test. Matlab's built-in functions (Breiman, 2001) were used. The Random Forest routine included 1200 trees, interaction curvature as the predictor selection, and the Gini diversity index as the splitting criterion. A cost matrix was used to compensate for the difference in sample size between the two OSA severity groups in the initial dataset (AHI ≥ 15 and AHI < 15). This procedure was repeated three times, thus resulting in three values for each of accuracy, sensitivity, and specificity for each feature combination. All values ≧0.7 were considered in the next step, and the difference between the maximum and minimum of each of the three values was evaluated. The maximum difference for each of accuracy, sensitivity, and specificity was recorded. For each feature combination, the mean values for each of accuracy, sensitivity, and specificity were recorded. The feature combination with a value between the maximum mean value and the difference between the maximum mean value and the maximum difference (or 2%) was selected as the best combination.
[0069] Using the best feature combinations selected in the previous step for each anthropometric subset, and using a Random Forest classifier with the same properties as described above, the classification accuracy, sensitivity, and specificity of the out-of-bag validation and blinded test datasets were evaluated. This process was repeated five times, and then the average values were evaluated. For each anthropometric subset, the feature combination with the highest validation and blinded test accuracy was selected as the best feature combination for that subset. The Random Forest classifier is preferable to the previously used SVM because: 1) respiratory sound signals are stochastic signals, and Random Forests are inherently random; and 2) OSA disorders have many confounding factors that affect respiratory sounds and make OSA disorders heterogeneous and complex. Therefore, a Random Forest classifier is preferably used to have multiple thresholds for each feature and to overcome complexity and heterogeneity.
[0070] For the final overall classification, the next step was to include and not include anthropometric features along with the acoustic features. Overall classification was evaluated using feature combinations, with each combination selected from the best ones for each subset. Since there was more than one final best feature combination per subset, various combinations were created. The combination that provided the greatest overall classification accuracy, sensitivity, and specificity for the out-of-bag validation and blinded test datasets was selected as the best feature combination. First, within each subset, each individual's classification decision was evaluated by assigning a label of 1 or -1 to each class (OSA group and non-OSA group). Using the results of the out-of-bag validation subset, the label 1 was multiplied by the sensitivity, and the label -1 was multiplied by the specificity. After performing the previous two steps for all subsets, the weighted classification decisions for each individual were averaged to arrive at the final classification decision. Any values >0 or <0 were classified into the OSA or non-OSA group, respectively.
[0071] Incorrectly classified individuals were examined for any commonalities between OSA subjects who were incorrectly classified as non-OSA and / or between non-OSA subjects who were incorrectly classified as OSA. The impact of ignoring subset results on the overall classification process was investigated. Correlation coefficients between AHI and anthropometric variables were examined. Subjects were classified using available variables from STOP-BANG (i.e., Bang; BMI, age, NC, and gender), and then classification accuracy, sensitivity, and specificity were evaluated. Correlation coefficients between AHI and the final selected acoustic features were examined. For each subset, the correlation coefficients between the feature combination and AHI, and between the feature combination and the logarithm of AHI, were evaluated using the final selected feature combination.
[0072] Experiment B Feature reduction and selection are important stages for data analysis, reducing training time and overfitting and facilitating feature interpretation. Most popular feature reduction techniques rely on selecting features highly correlated with the response and with minimal mutual information between them. In this experiment, an improved algorithm embodying all of the steps in Figures 3 through 6 was used to reduce features and select and model the most effective features for high-quality classification results. The results were compared with five existing popular feature reduction and selection techniques. The operability of the improved algorithm was demonstrated using a dataset adopted from a study of obstructive sleep apnea (113 participants as the training dataset and 86 participants as the blinded test dataset). Features were preprocessed and then modeled using three-, four-, and five-feature combinations. Models with high correlation with the response and low overlap percentages were selected for use in a random forest classification process. The models were then combined to provide higher accuracy. The classification accuracy resulting from the use of the improved algorithm was 25% higher than that resulting from the use of the five existing, popular feature reduction and selection techniques. In addition, the improved algorithm was approximately 20 times faster than the popular techniques. The improved technique can reduce features and select the best set for a high-performance classification process.
[0073] Features are the most important element in the classification process. Having irrelevant features leads to a useless classification process (failure). Features must be extracted so that they are relevant to the response. Therefore, by combining these features in a model, it is possible to reconstruct the response (regression process) or its main behavior (classification process). However, extracting many features leads to extensive training time, causes the curse of dimensionality, increases overfitting, and makes interpretation difficult for researchers and users. The improved algorithm reduces, selects, models, and utilizes features to provide better classification decisions.
[0074] Having many features can sometimes result in redundancy, which can negatively impact the generalization of the classification process. However, all relevant features carry information that may be important for a better classification process. However, as mentioned above, increasing the number of features is undesirable. Therefore, for a successful classification process, it is desirable to reduce the number of features while preserving all important information. Many studies and algorithms have been used to reduce the number of features by reducing feature redundancy and selecting features that are most relevant to the response (Battiti, 1994; Brown, 2009; Fleuret, 2004; Peng & Ding, 2005; Yang & Moody, 1999). This usually occurs by selecting features that are most relevant to the response and have the least mutual information. One prior example (Battiti, 1994) selects all features that provide the greatest mutual information with the response, while no selected features can be predicted by other selected features. Another example (Yang & Moody, 1999) selects features that together increase joint mutual information about the response and excludes any features that do not provide additional information, while another (Peng & Ding, 2005) selects features that are maximally dependent on the response and maximally independent of each other.
[0075] The improved algorithm disclosed herein processes features from the feature extraction stage to the final classification decision stage. This algorithm operates only on numerical, ordinal responses. The algorithm subjects the features to a modeling sequence (step 402) that reproduces multi-feature interaction effects on the responses. It then subjects the models to a high correlation and variance reduction stage with the responses (steps 403 to 410), followed by a classification step (step 412) and a model combination step (step 501). The utility of the improved algorithm is demonstrated through a practical example dataset. This dataset belongs to participants with and without obstructive sleep apnea (OSA), characterized by repeated periods of complete (apnea) or partial (hypopnea) cessation of breathing due to laryngeal collapse, and the apnea-hypopnea index (AHI) is a measure of the severity of OSA. The dataset for Experiment B was adopted from the dataset for Experiment A (i.e., 199 participants). Recordings were conducted approximately 1 to 2 hours before the polysomnography (PSG) study. For details on the recording protocol and preprocessing steps, see (Elwali & Moussavi, 2016). Study participants' PSG study reports were obtained after the overnight PSG measurement analysis was completed by a sleep technician. Table 1 (from above) shows the anthropometric statistics of the 199 study participants (83 women). The subject dataset was divided into two severity groups (AHI ≥ 15 and AHI < 15), training / validation (113 participants, 45 women), and blinded testing (86 participants, 38 women).
[0076] Twenty-seven features (anthropometric and acoustic features) were adopted from Experiment A. For more detailed information about the selection method of the adopted features, the adopted features from Experiment A had the following properties: 1) had a p-value of <0.05 between individuals with AHI ≥ 20 and AHI ≤ 10.2; 2) were normally distributed; 3) had low correlations with each other; and 4) outliers were removed.
[0077] Five existing, popular feature reduction selection techniques were used to reduce the features and select a set of 3, 4, or 5 of the 27 features used in the classification process. The five feature reduction techniques are Mutual Information Maximization (MIM) (Brown 2009), Maximum-Relevance Minimum-Redundancy (MRMR) (Peng et al., 2005), Mutual Information-Based Feature Selection (MIFS) (Battiti, 1994), Joint Mutual Information (JMI) (Yang & Moody, 1999), and Conditional Mutual Information Maximization (CMIM) (Fleuret, 2004). For these techniques, Matlab™2019 functions created by Stefan (2019) were used. The selected sets were used in the validation and blind testing phases. The method used in Experiment B, as already described above in connection with FIGS. 5 and 6 of the illustrated embodiment, concerns the nature of the responses and the three main stages—linearization and model generation, feature reduction and selection techniques, and model combination—and therefore will not be repeated here. Instead, attention is directed to the experimental results. The initial dataset statistics for this study were already described above in Table 1. The responses (AHI) had a high positive skewness (1.73). Therefore, the responses were scaled by taking their logarithms. The new value of the skewness was −0.041, as shown in FIG. 15. Directly reducing and selecting the feature sets using the five feature reduction techniques described above resulted in two 3-combinations, two 4-combinations, and four 5-combinations, with some combinations being repeated. Using these combinations in the classification process resulted in average validation and test classification accuracies of 65.6% and 68.4%, sensitivities of 57.4% and 59%, and specificities of 72.8% and 76.3%, respectively. See Table 5 for more detailed information.
[0078]
Table 5
[0079] When using the three-feature model, the average classification accuracy of the validation and blinded test datasets was 71.5% and 66.3% for the inventive method of the exemplified embodiment of the present invention, 63.8% and 61.6% for MIM, 63.8% and 63.2% for MRMR, 64.2% and 64.7% for MIFS, 64.6% and 63.1% for JMI, and 67.1% and 65.2% for CMIM. The most balanced (≥60% between validation and test, sensitivity and specificity) classification accuracy of the validation and blinded test datasets using the three-feature model was 73.2% and 73.8% for the inventive method, 74.3% and 74% for MIM, 71.3% and 74% for MRMR, 71.3% and 74% for MIFS, 72.3% and 74% for JMI, and 71.3% and 74% for CMIM. Using the five existing techniques above, the models selected for all of them yielded validation and blind classification accuracies of 74.3% and 74%, respectively. The highest results (for each of the five techniques) using the three-feature combination model are presented in Table 6, where the values in parentheses indicate the percentage of data used.
[0080] [Table 6-1]
[0081] [Table 6-2]
[0082] When using the four-feature model, the average classification accuracy for the validation and blinded test datasets was 75.7% and 69.6% for the method of the present invention, 68.6% and 65.6% for MIM, 71.2% and 67.1% for MRMR, 67.1% and 66.9% for MIFS, 71.6% and 65% for JMI, and 67.7% and 67.2% for CMIM, respectively. The classification accuracy for the validation and blinded test datasets using the most balanced (≥65% between validation and test, sensitivity and specificity) four-feature model was 80% and 77.3% for the method of the present invention, 78.9% and 72.2% for MIM, 78.9% and 73.4% for MRMR, 72.3% and 71.4% for MIFS, 74.3% and 71.4% for JMI, and 70% and 71.3% for CMIM, respectively. Using the five techniques above, the models selected for all of them yielded validation and blind classification accuracies of 67.4% and 71.5%, respectively. The highest results (for each of the five techniques) using the four-feature combination model are presented in Table 7, with the values in parentheses indicating the percentage of data used.
[0083] [Table 7-1]
[0084] [Table 7-2]
[0085] When using the five-feature model, the average classification accuracy for the validation and blinded test datasets was 81.2% and 71.3%, respectively, for the method of the present invention, 76.3% and 66.6% for MIM, 74.7% and 67.4% for MRMR, 74% and 63.8% for MIFS, 75.9% and 67.7% for JMI, and 75.3% and 67.9% for CMIM. The most balanced (>70% between validation and test, sensitivity and specificity) classification accuracy for the validation and blinded test datasets when using the five-feature model was 81.4% and 78.8%, respectively, for the method of the present invention, 78.9% and 75.6% for MIM, 79.8% and 75.6% for MRMR, 78% and 75.6% for MIFS, 82.4% and 74.7% for JMI, and 83.3% and 74.7% for CMIM. Using the five techniques above, the models selected for all of them yielded validation and blind classification accuracies of 63.4% and 50%, respectively. The highest results (for each of the five techniques) using the five-feature combination models are presented in Table 8, with the values in parentheses indicating the percentage of data used.
[0086] [Table 8-1]
[0087] [Table 8-2]
[0088] When using three features per model (2,925 models), the average time to perform each step to find the best 20 models was 8 seconds and 155.4 seconds for the inventive method and the popular reduction technique, respectively. When using four features per model (17,550 models), the average time to perform each step to find the best 20 models was 46.7 seconds and 937.2 seconds for the inventive method and the popular reduction technique, respectively. When using five features per model (80,730 models), the average time to perform each step to find the best 20 models was 269.6 seconds and 4,464 seconds for the inventive method and the popular reduction technique, respectively. Thus, the inventive method was ~20 times faster than the five popular methods.
[0089] The combined model improved classification results and covered more participants. When using a combination of three feature models, the maximum accuracy for the validation and blinded test datasets was 78.4% and 76.5%, respectively, covering 98.5% of the datasets. When using a combination of four feature models, the maximum accuracy for the validation and blinded test datasets was 79.3% and 81.1%, respectively, covering 98.5% of the datasets. When using a combination of five feature models, the maximum accuracy for the validation and blinded test datasets was 88.2% and 83.5%, respectively, covering 98% of the datasets. See Table 9 for the results.
[0090] [Table 9]
[0091] Every classification process depends on a set of features, and these features are related to the response. In many cases, the number of extracted features is huge, and some of these features are inefficient and redundant, which may cause overfitting. In addition, as the number of features increases, the computational cost also increases. Therefore, the process of reducing the number of features is an important step in order to find the best set of features in a short time, which can provide high classification accuracy and generalization sensitivity.
[0092] After extracting features and applying basic feature reduction techniques, many features may remain, resulting in a large feature set. Therefore, it is desirable to select the best set of features. In Experiment B, 182 features were extracted from the dataset, and after applying the basic feature reduction techniques described above, 27 different significant features resulted. In this case, it is desirable to select the best and final set of features (e.g., 3, 4, 5 features). However, trying all combinations (27C3, 4, 5) would take a lot of time. Therefore, it is desirable to adopt a novel feature combination selection method in a preferred embodiment of the inventive algorithm.
[0093] Using the five aforementioned feature selection techniques, which involve selecting features that are highly correlated with the response and have low redundancy among themselves, classification accuracy was poor for both the validation and blinded test datasets, even when using three, four, and five features in each combination. Furthermore, increasing the number of features did not improve the classification process. Therefore, directly using features in the modeling algorithm (random forest classifier) was not the best approach in the classification process. For this reason, a preferred embodiment of the disclosed invention models the response using available features and then uses a machine learning algorithm with the generated model to obtain the best classification decision. The generated model is believed to model the relationship between feature interactions and the response in a linear relationship. In addition, to have a better model for representing the response, the skewness of the response distribution, and therefore the above-mentioned scaling of the response, should be low. Referring to FIG. 16, the resulting model showed promising results for representing and predicting the response.
[0094] By scaling responses and building models, five feature reduction and selection techniques outperformed direct feature use, as seen in Tables 5, 6, 7, and 8. However, such methods were time-consuming, and classification rates were generally lower than those of the new techniques. In the inventive technique, all possible models were calculated, and then the one with the highest correlation to the response was selected. Here, no reliance was placed on the overall correlation, as this would have ignored any adverse correlations in the intermediate segments. This ensured linearity throughout the dataset. The disclosed algorithm only works on responses with severity detection (numerical or ordinal variables).
[0095] Having a high correlation does not ensure that there is no or low sparsity (misclassification) of participants near the midpoint between the actual and predicted responses. Therefore, an effort was made to find a model with minimal variance around the linearity between the response and prediction from the two perspectives. This was done by reducing the percentage of subject overlap between the segments from the response and the predicted perspective, which ensured the selection of a model with high linearity and low sparsity around the midpoint.
[0096] Furthermore, the numerical or ordinal variable technique also addressed the lack of data by weighting voting to involve most of the available participants in the dataset. The proposed method was more effective, accurate, and faster than using the five above-mentioned existing techniques, as shown in Table 9. This demonstrates promising results from the inventive algorithm, which can reduce features, provide better representation, reduce computational cost (between selection and classification), and improve classification accuracy.
[0097] Microphone Coupler The spectral characteristics of respiratory sounds provide a wealth of information regarding airflow and upper airway pathology. Tracheal and other respiratory sounds are commonly recorded using a microphone or an accelerometer attached to the surface of the skin [1, 2]. When using a microphone, the volume of air between the skin surface and the microphone's sensing element must convert upper airway or chest wall vibrations into measurable sound pressure [3]. A fixture known as a microphone coupler allows the microphone to be attached to the skin by creating a cavity (defined as an air chamber) in the microphone coupler, creating a sealed pocket of air between the microphone's sensing element and the skin surface. Proper air chamber design is important to ensure optimal signal reception by the microphone. New microphone hardware was developed based on a review of existing designs for the air chamber and coupler, and the results are summarized below.
[0098] A microphone coupler with an air chamber must be designed to accommodate the microphone being used, so that it is securely held and benefits from the effects of the air chamber. The applied design criteria for couplers with an air chamber are two-fold: 1) the coupler allows for easy and secure attachment to the skin and also serves to properly position the sensing element of the microphone with the base of the air chamber, and 2) the effect of the air chamber benefits the signal being recorded by allowing the full frequency range of the microphone to be heard, while having the desired effect of optimal sensitivity over a relatively wide (>5 kHz) frequency band of sound collection.
[0099] The novel coupler was custom designed and manufactured from a single, solid piece of ultra-high molecular weight polyethylene (UHMWPE). An air chamber was cut into the coupler's front surface, and a counterbore was cut on the rear side to properly mount and position the microphone at the base of the air chamber. In addition, the front flat surface surrounding the air chamber was designed to accommodate a two-sided adhesive disk (Manufacturer: 3M, Product Number: 2181) that provides significant adhesion for secure attachment to the skin.
[0100] The geometry of the air chamber was chosen to be conical because it has been proven that a conical air chamber provides 5-10 dB higher sensitivity than a cylindrical air chamber of the same diameter.[9] The design of a conical air chamber is then divided into three fundamental dimensional characteristics: (i) the dimensions of the base of the air chamber (i.e., the microphone capsule), (ii) the dimensions of the top of the air chamber (i.e., the surface in contact with the skin), and (iii) the depth of the air chamber.
[0101] The diameter of the air chamber base was chosen to be 4.85 mm. This dimension was chosen to allow for full, unobstructed exposure of the entry port to the microphone. The design of the air chamber top diameter and depth was established using literature recommendations, stating that overall high-frequency response decreases with deeper air chambers [3], with an optimal air chamber opening diameter of 10 to 15 mm [9]. Air chambers were fabricated with top diameters of 10.00 mm (Figure 17), 12.50 mm, and 15.00 mm. The air chamber depth was chosen to be 2.60 mm.
[0102] The three microphone coupler designs described above placed the microphone capsule in direct contact with the coupler's UHMWPE material. A second modification of our microphone coupler design incorporates a slightly larger counterbore to accommodate a vibration-isolating rubber ring (1.50 mm thick) surrounding the microphone capsule. With the vibration-isolating ring, the metal body of the microphone capsule is no longer in direct contact with the coupler's UHMWPE material. With a particular focus on examining higher frequency components, the addition of the vibration-isolating ring was chosen for testing, as the chamber exhibited optimal responses between approximately 150 and 10,000 Hz for inspiration and 130 and 8,500 Hz for expiration. The same signal was applied to the air chamber with and without the vibration-isolating ring, and the responses were compared.
[0103] The AOM-5024L-HD-R electret condenser microphone was used because of its high fidelity in the frequency range of breathing sounds. This microphone was used to test various designs of the air chamber and coupler described herein. The microphone selected had cylindrical dimensions of 9.7 mm diameter and 5.0 mm height, an 80 dB signal-to-noise ratio (1 kHz, 94 dB input, A-weighted), a frequency range of 20–20,000 Hz, and omnidirectional characteristics.
[0104] The microphone was connected using the recommended driver circuit described in the datasheet, with component values of 2.2 kΩ and 0.1 ΩF for resistance and capacitance, respectively. Digitalization of the analog signal from the microphone was performed using custom-designed hardware consisting of an SGTL5000 stereo codec (Fe-Pi Audio Z v2), an ARM1176JZF-S processor (Raspberry Pi Zero W), and a custom-made precision-regulated power supply. The audio information was losslessly saved in a WAV-style file format with a sampling rate of 44,100 Hz and 16-bit accuracy. Audio recordings were made by attaching the microphone coupler (with the attached microphone) with a double-sided adhesive ring. For each recording, the air chamber and microphone were positioned in the suprasternal recess of the trachea after cleaning with isopropyl alcohol. A new adhesive ring was used each time the coupler was replaced.
[0105] All recordings were completed in an anechoic chamber (>30 dB noise reduction). The same recording device hardware was used on the same human subject for each recording. For each recording, subjects were guided by the investigator to breathe only through the mouth, starting with inspiration and following a breathing pattern of five complete breathing cycles at normal flow rate. At the end of each cycle, subjects were instructed to hold their breath for five seconds.
[0106] The recorded data were preprocessed and normalized using procedures described in previous work [6, 8]. First, each recording was manually inspected in the time-frequency domain to eliminate artifacts, speech noise, tones, interruptions, or respiratory phases containing very low signal-to-noise ratios compared to background noise (assessed during breath-hold segments). Next, respiratory sounds were filtered using a fourth-order bandpass (75 to 20,000 Hz) filter (Butterworth). Finally, each respiratory phase signal was normalized to its energy to remove the effects of plausible airflow fluctuations during the respiratory cycle [8]. The power spectrum of the normalized data was then estimated using the Welch method with a window size of 1,024 samples (~23.22 ms) and a 50% overlap between adjacent windows. The average power spectrum was evaluated for each respiratory phase (inspiration and expiration) and varied the diameter of the air chamber opening.
[0107] During inspiration (Figure 18), the 15.00 mm chamber provided a better (higher amplitude) response below approximately 150 Hz. Between approximately 150 and 430 Hz, the 10.00 mm chamber was observed to provide the optimal response, but for a small range from approximately 430 to 640 Hz, the 12.50 mm chamber dominated. From 640 to 815 Hz, the 12.50 mm chamber provided a slightly better response than the 10.00 mm chamber, which showed roughly equivalent responses. From 815 Hz to approximately 3750 Hz, the 10.00 mm and 12.50 mm chambers continued to show similar responses, with both clearly outperforming the 15.00 mm chamber. From approximately 3750 to 10,000 Hz, the 10.00 mm chamber provided significantly better response than both the 12.50 mm and 15.00 mm chambers. Above 10 kHz, the 15.00 mm chamber showed the best response, but in this frequency range the signal was low amplitude and noisy.
[0108] For the expiratory signal (Figure 19), many of the same trends were seen from the inspiratory signal, with slight differences: below approximately 130 Hz, the 15.00 mm provided the best response of the three chambers. In the ranges between approximately 130 and 220 Hz, 300 and 340 Hz, 300 and 340 Hz, and 450 and 580 Hz, the 12.50 mm chamber provided the best response, while the 15.00 mm chamber showed a better response in a small range between approximately 220 and 300 Hz, with the 10.00 mm chamber having a dominant response between approximately 340 and 450 Hz. From 580 to 2,500 Hz, the 10.00 mm chamber provided the best response, with the 15.00 mm chamber dominating in the small range from 2,500 to 2,800 Hz, after which the 10.00 mm chamber had a significantly better response up to about 3,700 Hz, where the response curves for the 10.00 mm and 15.00 mm chambers became very similar. After about 8,500 Hz, a noticeable change was observed, where both the 12.50 mm and 15.00 mm chambers provided better response compared to the 10.00 mm chamber, but in this range the signal amplitude was again very low and noisy.
[0109] Figure 20 shows the response curves of a 10.00 mm air chamber with and without the addition of a rubber vibration isolation ring around the microphone during inspiration and expiration. During inspiration, from approximately 40 to 200 Hz, the air chamber without the rubber vibration isolation ring provided a better response. Between approximately 200 and 890 Hz, the rubber vibration isolation ring provided a better response, but above 890 Hz, the chamber without the rubber vibration isolation ring provided a slightly better response. During expiration, below approximately 430 Hz, the air chamber without the rubber vibration isolation ring provided a better response. Between approximately 430 and 820 Hz, the rubber vibration isolation ring provided a better response, but above 820 Hz, a better response was observed without the rubber vibration isolation ring.
[0110] The power spectra of the inspiratory and expiratory respiratory phases showed clear differences as a function of frequency for the optimal air chamber: <150 Hz, 150 to 10,000 Hz, and >10,000 Hz for inspiration, and <130 Hz, 130 to 8,500 Hz, and >8,500 Hz for expiration. These differences can be explained by several physical properties of the air chamber.
[0111] First, acoustic resonance as a function of air chamber geometry acts to increase sound pressure at frequencies close to the natural frequency of each air chamber. Calculations show that the effect of resonance can amplify sound pressure by approximately three times for frequencies up to near 0 Hz, with this effect exponentially decaying until approximately 150 Hz, after which resonance has little or no effect. Spectra acquired at frequencies <150 Hz and <130 Hz (during inspiration and expiration, respectively) were attributed to the effect of resonance as a function of air chamber geometry.
[0112] Second, air chambers with larger opening diameters present an increased surface area that acts to absorb sound over a larger range, thereby reducing the sound pressure reaching the microphone. Calculations show that the acoustic impedance of the smallest-opening air chamber is twice as large as that of the largest-opening air chamber. The higher the chamber's acoustic impedance, the less sound is absorbed by the material. Because the effect of resonance is not seen at frequencies >150 Hz and >130 Hz (in the inhalation and exhalation phases, respectively), the influence of acoustic impedance as a function of air chamber surface area can explain the results seen in the intermediate ranges, i.e., between 150 and 10,000 Hz for inhalation and between 130 and 8,500 Hz for exhalation.
[0113] Conversely, larger opening air chambers have the advantage of capturing a larger area of surface vibration (at the skin), which has the overall effect of capturing a greater range of sound compared to smaller opening air chambers. In the frequency ranges above 10,000 Hz and 8,5000 Hz (in the inhalation and exhalation phases, respectively), this increased sound capture appears to outweigh the effect of acoustic impedance, making the largest opening air chamber the best choice.
[0114] Finally, since mechanical vibrations on the microphone capsule can be translated into vibrations in the microphone diaphragm, the anti-vibration rubber ring provides a benefit independent of the nature of the air chamber by absorbing any vibrations transmitted through the material of the microphone coupler.
[0115] In general, air chambers with smaller openings performed better at frequencies ranging from 150 to 10,000 Hz for inspiration and from 130 to 8,500 Hz for expiration. Above and below these ranges, chambers with the largest openings provided optimal sensitivity. The addition of a vibration-isolating rubber ring proved beneficial in increasing the sensitivity of the 10.00 mm chamber in the frequency ranges of approximately 200 to 890 Hz for inspiration and 430 to 820 Hz for expiration. Given that many existing recording systems cited in the literature have an upper limit of 5 kHz, the use of an air chamber with a smaller opening diameter is recommended unless the focus is on analyzing frequencies below approximately 150 Hz or 130 Hz (for inspiration and expiration, respectively). Similarly, the use of vibration-isolating rubber may prove beneficial, depending on the frequency range studied.
[0116] In the context of the present invention, a particular microphone air chamber optimized for use in OSA screening has been described throughout the above, but it should be understood that other microphone designs may alternatively be used. Similarly, various modifications may be made to the preferred embodiments described hereinabove, thereby obviously resulting in many widely varying embodiments, and all matters including this specification should be interpreted as illustrative only and not limiting.
[0117] literature OSA upon waking - Experiment A Ahbab, S., Ataoglu, HE, Tuna, M., Karasulu, L., Cetin, F., Temiz, LU, & Yenigun, M. (2013). Neck circumference, metabolic syndrome and obstructive sleep apnea syndrome; evaluation of possible linkage. Medical Science Monitor : International Medical Journal of Experimental and Clinical Research, 19, 111-117. doi:10.12659 / MSM.883776 [doi] American Academy of Sleep Medicine. (2005). International classification of sleep disorders: Diagnostic and coding manual (2nd ed.). Westchester: American Academy of Sleep Medicine. American Society of Anesthesiologists. (2006). Practice guidelines for the perioperative management of patients with obstructive sleep apnea: A report by the american society of anesthesiologists task force on perioperative management of patients with obstructive sleep apnea. Anesthesiology, 104(5), 1081. doi:00000542-200605000-00026 [pii] Barsties, B., Verfaillie, R., Roy, N., & Maryn, Y. (2013). Do body mass index and fat volume influence vocal quality, phonatory range, and aerodynamics in females? Paper presented at the CoDAS, , 25(4) 310-318. Benjamini, Y. (1988). Opening the box of a boxplot. The American Statistician, 42(4), 257-262. Bixler, E. O., Vgontzas, A. N., Ten Have, T., Tyson, K., & Kales, A. (1998). Effects of age on sleep apnea in men: I. prevalence and severity. American Journal of Respiratory and Critical Care Medicine, 157(1), 144-148. Bonsignore, M. R., Baiamonte, P., Mazzuca, E., Castrogiovanni, A., & Marrone, O. (2019). Obstructive sleep apnea and comorbidities: A dangerous liaison. Multidisciplinary Respiratory Medicine, 14(1), 8. Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5-32. Chung, F., & Elsaid, H. (2009). Screening for obstructive sleep apnea before surgery: Why is it important? Current Opinion in Anaesthesiology, 22(3), 405-411. doi:10.1097 / ACO.0b013e32832a96e2 [doi] El-Sayed, I. H. (2012). Comparison of four sleep questionnaires for screening obstructive sleep apnea. Egyptian Journal of Chest Diseases and Tuberculosis, 61(4), 433-441. Elwali, A., & Moussavi, Z. (2017). Obstructive sleep apnea screening and airway structure characterization during wakefulness using tracheal breathing sounds. Annals of Biomedical Engineering, 45(3), 839-850. Epstein, L. J., Kristo, D., Strollo, P. J.,Jr, Friedman, N., Malhotra, A., Patil, S. P., . . . Adult Obstructive Sleep Apnea Task Force of the American Academy of Sleep Medicine. (2009). Clinical guideline for the evaluation, management and long-term care of obstructive sleep apnea in adults. Journal of Clinical Sleep Medicine : JCSM : Official Publication of the American Academy of Sleep Medicine, 5(3), 263-276. Finkelstein, Y., Wolf, L., Nachmani, A., Lipowezky, U., Rub, M., & Berger, G. (2014). Velopharyngeal anatomy in patients with obstructive sleep apnea versus normal subjects. Journal of Oral and Maxillofacial Surgery, 72(7), 1350-1372. Glass, G. V., Smith, M. L., & McGaw, B. (1981). Meta-analysis in social research Sage Publications, Incorporated. Goldshtein, E., Tarasiuk, A., & Zigel, Y. (2011). Automatic detection of obstructive sleep apnea using speech signals. Biomedical Engineering, IEEE Transactions On, 58(5), 1373-1382. Gross, J. B., Apfelbaum, J. L., Caplan, R. A., Connis, R. T., Cote, C. J., Nickinovich, D. G., . . . Ydens, L. (2014). Practice guidelines for the perioperative management of patients with obstructive sleep apnea an updated report by the american society of anesthesiologists task force on perioperative management of patients with obstructive sleep apnea. Anesthesiology, 120(2), 268-286. Gupta, S., Sharma, R., & Jain, D. (2005). Airway assessment: Predictors of difficult airway. Indian J Anaesth, 49(4), 257-262. Higuchi, T. (1988). Approach to an irregular time series on the basis of the fractal theory. Physica D: Nonlinear Phenomena, 31(2), 277-283. Hurst, H. E. (1951). Long-term storage capacity of reservoirs. Trans.Amer.Soc.Civil Eng., 116, 770-808. Jung, D. G., Cho, H. Y., Grunstein, R. R., & Yee, B. (2004). Predictive value of kushida index and acoustic pharyngometry for the evaluation of upper airway in subjects with or without obstructive sleep apnea. Journal of Korean Medical Science, 19(5), 662-667. doi:200410662 [pii] Karimi, D. (2012). Spectral and bispectral analysis of awake breathing sounds for obstructive sleep apnea diagnosis (MSc). Katz, M. J. (1988). Fractals and the analysis of waveforms. Computers in Biology and Medicine, 18(3), 145-156. Kriboy, M., Tarasiuk, A., & Zigel, Y. (2014). Detection of obstructive sleep apnea in awake subjects by exploiting body posture effects on the speech signal. Paper presented at the Engineering in Medicine and Biology Society (EMBC), 2014 36th Annual International Conference of the IEEE, 4224-4227. Lan, Z., Itoi, A., Takashima, M., Oda, M., & Tomoda, K. (2006). Difference of pharyngeal morphology and mechanical property between OSAHS patients and normal subjects. Auris Nasus Larynx, 33(4), 433-439. Lilliefors, H. W. (1967). On the kolmogorov-smirnov test for normality with mean and variance unknown. Journal of the American Statistical Association, 62(318), 399-402. Linville, S. E. (2001). Vocal aging Singular Thomson Learning. Littner, M. R. (2007). Mild obstructive sleep apnea syndrome should not be treated. con. Journal of Clinical Sleep Medicine : JCSM : Official Publication of the American Academy of Sleep Medicine, 3(3), 263-264. Medscape. (Aug 23, 2016, Should we treat mild obstructive sleep apnea? Montazeri, A., Giannouli, E., & Moussavi, Z. (2012). Assessment of obstructive sleep apnea and its severity during wakefulness. Annals of Biomedical Engineering, 40(4), 916-924. Moussavi, Z., Elwali, A., Soltanzadeh, R., MacGregor, C., & Lithgow, B. (2015). Breathing sounds characteristics correlate with structural changes of upper airway due to obstructive sleep apnea. Paper presented at the Engineering in Medicine and Biology Society (EMBC), 2015 37th Annual International Conference of the IEEE, 5956-5959. Nagappa, M., Liao, P., Wong, J., Auckley, D., Ramachandran, S. K., Memtsoudis, S., . . . Chung, F. (2015). Validation of the STOP-BANG questionnaire as a screening tool for obstructive sleep apnea among different populations: A systematic review and meta-analysis. PLoS One, 10(12), e0143697. Nikias, C. L., & Raghuveer, M. R. (1987). Bispectrum estimation: A digital signal processing framework. Proceedings of the IEEE, 75(7), 869-891. Proakis, J. (2006). DG monolakis-digital signal processing., 911. ResMed. (2013). Sleep apnea facts and figures. ResMed Corp. Sola-Soler, J., Fiz, J. A., Torres, A., & Jane, R. (2014). Identification of obstructive sleep apnea patients from tracheal breath sound analysis during wakefulness in polysomnographic studies. Paper presented at the Engineering in Medicine and Biology Society (EMBC), 2014 36th Annual International Conference of the IEEE, 4232-4235. Titze, I. (2000). Principles of voice production (national center for voice and speech, iowa city, IA). Chap, 6, 149-184. Torre III, P., & Barlow, J. A. (2009). Age-related changes in acoustic characteristics of adult speech. Journal of Communication Disorders, 42(5), 324-333. Tukey, J. W. (1977). Exploratory data analysis Reading, Mass. Vavougios, G. D., Natsios, G., Pastaka, C., Zarogiannis, S. G., & Gourgoulianis, K. I. (2016). Phenotypes of comorbidity in OSAS patients: Combining categorical principal component analysis with cluster analysis. Journal of Sleep Research, 25(1), 31-38. Young, T., Finn, L., Peppard, P. E., Szklo-Coxe, M., Austin, D., Nieto, F. J.,... Hla, K. M. (2008). Sleep disordered breathing and mortality: Eighteen-year follow-up of the Wisconsin sleep cohort. Sleep, 31(8), 1071-1078. OSA-Experiment B at waking up Battiti, R. (1994). Using mutual information for selecting features in supervised neural net learning. IEEE Transactions on Neural Networks, 5(4), 537-550. Brown, G. (2009). A new perspective for information theoretic feature selection. Paper presented at the Artificial Intelligence and Statistics, 49-56. Chatterjee, S., & Hadi, A. S. (1986). Influential observations, high leverage points, and outliers in linear regression. Statistical Science, 1(3), 379-393. Elwali, A., & Moussavi, Z. (2016). Obstructive sleep apnea screening and airway structure characterization during wakefulness using tracheal breathing sounds. Annals of Biomedical Engineering,, 1-12. breathing sounds. Scientific Reports, 9(1), 1-12. Fleuret, F. (2004). Fast binary feature selection with conditional mutual information. Journal of Machine Learning Research, 5(Nov), 1531-1555. Peng, H., & Ding, C. (2005). Minimum redundancy and maximum relevance feature selection and recent advances in cancer classification. Feature Selection for Data Mining, 52 Spearman, C. (1904). The proof and measurement of association between two things. American Journal of Psychology, 15(1), 72-101. Yang, H., & Moody, J. (1999). Feature selection based on joint mutual information. Paper presented at the Proceedings of International ICSC Symposium on Advances in Intelligent Data Analysis, 22-25. マイクロフォンカプラ H. Pasterkamp, S. S. Kraman, P. D. DeFrain and G. R. Wodicka, “Measurement of Respiratory Acoustical Signals: Comparison of Sensors,“Chest, vol. 104, no. 5, pp. 1518-1525, 1993. A. Yadollahi and Z. Moussavi, “Acoustical Flow Estimation: Review and Validation,“ IEEE Magazine in Biomedical Engineering, vol. 26, no. 1, pp. 56-61, 2007. G. R. Wodicka, S. S. Kraman, G. M. Zenk and H. Pasterkamp, “Measurement of Respiratory Acoustic Signals: Effect of Microphone Air Cavity Depth,“ Chest, vol. 106, pp. 1140-1144, 1994. A. Yadollahi, E. Giannouli and Z. Moussavi, “Sleep apnea monitoring and diagnosis based on pulse oximetery and tracheal sound signals,“ Medical & biological engineering & computing, vol. 48, no. 11, pp. 1087-1097, 2010. A. Yadollahi and Z. Moussavi, “The effect of anthropometric variations on acoustical flow estimation: proposing a novel approach for flow estimation without the need for individual calibration,“ IEEE Transactions on Biomedical Engineering, vol. 58, no. 6, pp. 1663-1670, 2011. A. Montazeri, E. Giannouli and Z. Moussavi, “Assessment of obstructive sleep apnea and its severity during wakefulness,“ Annals of biomedical engineering, vol. 40, no. 4, pp. 916-924, 2012. A. Elwali and Z. Moussavi, “A novel diagnostic decision making procedure for screening obstructive sleep apnea using anthropometric information and tracheal breathing sounds during wakefulness,“ Scientific Reports (Nature), vol. 9, no. 1, pp. 1-12, 2019. A. Elwali and Z. Moussavi, “Obstructive Sleep Apnea Screening and Airway Structure Characterization During Wakefulness Using Tracheal Breathing Sounds,“ Annals of Biomedical Engineering, vol. 45, no. 3, pp. 839-850, 2017. S. S. Kraman, H. Pasterkamp, G. R. Wodicka and Y. Oh, “Measurement of respiratory acoustic signals: effect of microphone air cavity width, shape and venting,“ Chest, vol. 108, no. 4, p. 1004-1008, 1995.
Claims
1. 1. A method of deriving a screening tool for obstructive sleep apnea (OSA), the method comprising: (a) obtaining an initial dataset comprising a respective subject dataset for each of a plurality of human test subjects whose respective audio recordings were taken during a period of fully awake state, said subject dataset comprising at least: the test subject's known OSA severity parameters; anthropometric data identifying different anthropometric parameters of the test subject; and audio data including stored audio signals from said audio recordings of each of said test subjects; and a process in which (b) extracting at least spectral and bispectral features from the speech data of the subject data set; (c) selecting a training dataset from the initial dataset, and from the training dataset, grouping together subject datasets representing a first high severity group of the test subjects whose known OSA severity parameter is above a first threshold, and grouping together subject datasets representing a second low severity group of the test subjects whose known OSA severity parameter is below a second threshold that is less than the first threshold; (d) dividing the subject data set for each of the high severity group and the low severity group into a plurality of anthropometrically distinct subsets based on the anthropometric data; (e) at least in part, for each anthropometrically distinct subset, filtering the extracted features into a selected feature subset; and Creating a respective pool of predictive models for each anthropometrically distinct subset that predict OSA severity values. deriving inputs for a classifier training procedure by: (f) training a classifier using the derived inputs to classify human patients as either OSA or non-OSA based on a respective patient dataset, the patient dataset comprising at least: anthropometric data identifying different anthropometric parameters of the patient; and audio data comprising audio signals stored from audio recordings of each of said patients during a period of full wakefulness; and / or at least one of the feature data relating to features already extracted from the audio signals from the audio recordings of each of the patients; and A method comprising:
2. (g) in non-transitory computer-readable memory, the trained classifier; and for each anthropometrically distinct subset, identifying a respective appropriate feature set for classification of patient anthropometric parameters that overlap with anthropometric parameters of the subject dataset of the anthropometrically distinct subset; statements and instructions executable by one or more computer processors; and further comprising the step of storing the The said sentences and commands are: reading the anthropometric data of the patient data set; comparing the anthropometric data with each appropriate feature set; selecting which particular features are required from the speech data or the feature data of the patient data set based on the comparison; 10. The method of claim 1, wherein the specific features are input into the trained classifier, executable to classify the patient as either OSA or non-OSA.
3. 3. The method of claim 1 or 2, wherein step (e) further comprises dividing each data point of the predictive model in each pool into a first set of different OSA severity groups based on the actual OSA severity value from the subject dataset for the test subject.
4. 4. The method of claim 3, wherein step (e) further comprises calculating a respective correlation value for each predictive model in each pool and removing from each pool any predictive model whose absolute value of the correlation value is less than a correlation threshold.
5. Each of the correlation values is a population correlation value, and the population correlation value is Calculating the correlation for each severity group in the model; calculating an average of the correlation values of said group; and the average of the group correlation values is then used as the population correlation value.
6. The method according to claim 4 or 5, wherein the correlation threshold is calculated as a predetermined percentage of the absolute value of the largest correlation value among the respective correlation values.
7. Step (e) comprises, for each of the predictive models in each pool: calculating a mean OSA severity value for the first set of different OSA severity groups; determining whether the mean OSA severity value increases or decreases upon matching sequences with the first set of different OSA severity groups; and removing from the pool any predictive models that do not increase or decrease the mean OSA severity value when matching sequences with the first set of OSA severity groups. The method of any one of claims 3 to 6, further comprising:
8. 8. The method of claim 7, wherein the mean OSA severity value is calculated as the average of the predicted OSA severity values from the model.
9. Step (e) comprises, for each of the predictive models in each pool: dividing the data points for each of the predictive models in the pool into a second set of different OSA severity groups based on the predicted OSA severity value from the model; calculating an average actual OSA severity value for the second set of different OSA severity groups; determining whether the average actual OSA severity value increases or decreases upon matching the sequence with the second set of OSA severity groups; and removing from the pool any predictive models that do not increase or decrease the mean actual OSA severity value when matching sequences with the second set of OSA severity groups. The method of claim 8 further comprising:
10. 10. The method of claim 3, wherein step (e) further comprises calculating, for each model in each pool, an overlap rate between adjacent OSA severity groups in each set of different OSA severity groups.
11. 11. The method of claim 10, wherein step (e) further comprises calculating an average of the overlap rates of the models for each set of different OSA severity groups for each model in each pool.
12. Step (e) is calculating, for each model in each pool, the overlap rate between adjacent OSA severity groups in both the first and second sets of different OSA severity groups; For each model in each pool, calculating the average overlap rate of the model in both the first set and the second set of different OSA severity groups.
10. The method of claim 9, further comprising:
13. 13. The method of claim 11 or 12, wherein step (e) further comprises using the average of the overlap rates for each model to filter each pool into a subset of models in each pool that have a smaller average overlap rate between the adjacent OSA severity groups than other models from the pool.
14. Step (e) is For each model in each pool, calculating an overall average of the average overlap rates from both the first and second sets of different OSA severity groups; and filtering each pool to a subset of models in each pool that have a low overall average overlap relative to other models from said pool; The method of claim 12 further comprising:
15. 15. The method of claim 1, wherein step (c) also comprises selecting a blind dataset from the initial dataset, and step (e) further comprises: running each model from each pool using the subject dataset from the blind dataset to calculate an estimated OSA severity value for the blind dataset; and assessing the accuracy of the estimated OSA severity value against actual OSA severity values from the blind dataset.
16. 16. The method of claim 15, wherein step (f) comprises, for each model in each pool, at least one training procedure having a classifier training step and a classifier testing step, wherein the classifier training step comprises running the classifier using model-predicted OSA severity values from a subset of the training dataset, and the classifier testing step comprises re-running the classifier using model-predicted OSA severity values from the same or a different subset of the training dataset, and evaluating the classification results obtained by re-running the classifier against actual OSA severity values from the same or the different subset of the training dataset.
17. 15. The method of claim 1, wherein step (c) also comprises selecting a blind dataset from the initial dataset; and step (f) comprises, for each model in each pool, at least one training procedure having a classifier training step and a classifier validation step, wherein the classifier training step comprises running the classifier using model-predicted OSA severity values from a subset of the training dataset; and the classifier validation step comprises re-running the classifier using model-predicted OSA severity values from the same or a different subset of the training dataset, and comparing the classification results obtained by re-running the classifier against actual OSA severity values from the same or a different subset of the training dataset to evaluate the performance of the model.
18. 18. The method of claim 16 or 17, wherein step (f) further comprises filtering each pool of models at least in part by comparing results obtained from a classifier testing step performed on models in the pool and removing models from the pool that have poor evaluated performance relative to other models in the pool.
19. Step (f) comprises, at least in part, Comparing the estimated accuracy of the models in the pool; and comparing the results obtained from the classifier testing step performed on the models of the pool; and removing models from the pool that have low evaluated accuracy and evaluated performance relative to other models in the pool; 17. The method of claim 16, further comprising filtering each pool of models by:
20. Step (f) is grouping different models together into model combinations, each model combination including multiple models from different pools; For each model combination, running the different models using at least a portion of the subject data sets of the initial data set and sending the OSA severity values predicted by the models to the classifier to obtain a respective classification result for each of the different models for each subject; For each subject whose subject data set has been run through the different models of the model combination, assigning one of two possible weighting factors to the classification result of each different model depending on whether the classification result is an OSA classification or a non-OSA classification; calculating, for each subject, a weighted average of the classification results derived from the different models of the model combination using the assigned weighting coefficients; For each subject, comparing the weighted average of the classification results of the model combinations against the actual OSA severity values from each of the subject data sets to assess whether the subject was correctly or incorrectly classified by each model combination, and based thereon, calculating one or more performance metrics for each model combination, including at least a respective overall classification accuracy based on how many subjects were correctly classified out of the total number of subjects classified by the model combination; selecting a best model combination based on a comparison of the one or more performance metrics calculated for the model combinations; The method of any one of claims 1 to 19, further comprising:
21. 21. The method of claim 20, wherein the two possible weighting factors assigned to each respective classification result are (a) the calculated sensitivity of the model assigned if the subject is classified as OSA, and (b) the calculated specificity of the model assigned if the subject is classified as non-OSA.
22. 22. The method of claim 20 or 21, wherein the calculated performance metrics for each model combination further include an overall classification sensitivity and an overall classification specificity of the model combination.
23. 23. The method of any one of claims 20 to 22, comprising after step (f) storing the best model combination together with the trained classifier in one or more computer readable memories for subsequent use in the screening tool.
24. 24. The method of any one of claims 1 to 23, wherein at least steps (a) through (f) are performed through the execution, by one or more computer processors, of statements and instructions stored on one or more non-transitory computer-readable media.
25. One or more non-transitory computer-readable media having executable statements and instructions stored thereon for performing at least steps (a) through (f) of any one of claims 1 to 24.
26. 26. One or more non-transitory computer-readable media according to claim 25, also having stored thereon the initial data set according to claim 1.
27. 27. The one or more non-transitory computer-readable media of claim 25 or 26, wherein the executable statements and instructions are further configured to extract the features from the audio data.
28. 28. The one or more non-transitory computer-readable media of any one of claims 25-27, wherein the executable statements and instructions, when executed by one or more processors of a computer to which a microphone is operably connected, are further configured to capture the audio recording through the microphone and store the audio data.
29. 1. A method of operating a system for screening a patient for obstructive sleep apnea (OSA), the method comprising: (a) A trained classifier derived in the manner as recited in steps (a) to (f) of any one of claims 1 to 24. a patient data set for said patient of the type listed in step (f) of claim 1; and statements and instructions executable by one or more computer processors acquiring one or more computer readable media having stored thereon: (b) through execution of said statements and instructions by said one or more computer processors; (i) reading the anthropometric data of the patient data set; (ii) running the trained classifier multiple times, each time starting with inputs consisting of or derived from different combinations of suitable features consisting of at least spectral features and bispectral features selected or derived from the patient dataset, in particular for different anthropometric parameters read from the anthropometric data of the patient dataset, and deriving therefrom, for each run of the trained classifier, a respective classification result classifying the patient as either OSA or non-OSA; (iii) deriving a final classification result for the patient based on the classification result of step (b)(ii); and (iv) storing or displaying the final classification result; A method comprising:
30. 30. The method of claim 29, wherein in step (b)(ii), the different combinations of features are initially selected by comparing the anthropometric data of the patient dataset against stored records identifying different respective appropriate feature sets of a plurality of different stored models that provide input data to the trained classifier.
31. 31. The method of claim 30, wherein step (b)(ii) comprises, after selecting the different combinations of features, inputting values for each combination of features into each of the different models to derive estimated OSA severity values from the models, which are then input into the trained classifier from which the classification result is derived.
32. One or more non-transitory computer-readable media having executable statements and instructions stored thereon for performing step (b) of any one of claims 29-31.
33. 33. The one or more non-transitory computer-readable media of claim 32, also having stored thereon the patient datasets listed in step (a) of claim 29.
34. 34. The one or more non-transitory computer-readable media of claim 32 or 33, wherein the executable statements and instructions are further configured to extract features from the audio data.
35. 35. The one or more non-transitory computer-readable media of any one of claims 32-34, wherein the executable statements and instructions, when executed by one or more processors of a computer to which a microphone is operably connected, are further configured to capture the audio recording through the microphone and store the audio data on the one or more non-transitory computer-readable media.
36. 1. A system for deriving or operating a screening tool for obstructive sleep apnea, the system comprising: one or more computers; the one or more computers one or more computer processors; one or more non-transitory computer-readable media coupled to the one or more computer processors; and a microphone connected to one of the one or more computers for capturing an audio recording of each of the human test subjects and / or patients during a period of full wakefulness and storing the audio recording as audio data on the one or more non-transitory computer-readable media; the one or more non-transitory computer-readable media also storing executable statements and instructions configured, when executed by the one or more computer processors, to extract features from the audio recording and perform at least steps (a) to (f) of any one of claims 1 to 24 and / or step (b) of any one of claims 29 to 31.
37. 25. The method of any one of claims 1 to 24, comprising first capturing an audio recording of each of the human test subjects using a microphone, storing the audio recording as audio data on a computer readable medium, and then computer-assisted or computer-automated extraction of the features from the audio data.
38. 32. The method of any one of claims 29 to 31, comprising first capturing the audio recording of each of the human patients using a microphone and storing the audio recording as audio data on a computer readable medium, and then computer-assisted or computer-automated extraction or selection of the appropriate features from the audio data or feature data.
39. Filtering the features extracted in step (e) includes: (i) the calculated statistical significance of each extracted spectral feature for the low severity group and the high severity group; (ii) the calculated normality of each feature extracted from randomly selected smaller subgroups within the high severity group and the low severity group that are identical to each other; and / or The method of any one of claims 1 to 24, based at least in part on one or more of (iii) calculated correlation coefficients between pairs of extracted features.
Citation Information
Patent Citations
Artificial Neural Network for Predicting Respiratory Disorders and Method for Developing It
JP2002505892A
A system and method for characterizing the upper airway using vocal characteristics.
JP2014532448A
Acoustic system and methodology for identifying the risk of obstructive sleep apnea during wakefulness
US20140142452A1