Capturing truncated proteoforms in exhaled breath for disease diagnosis and treatment
A packed-bed column system with mass spectrometry analysis addresses the challenge of noninvasive RTI diagnosis by capturing and analyzing cleaved proteoforms in exhaled aerosols, offering rapid and accurate RTI prediction.
Patent Information
- Application Number
- JP2025507548
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-11
- Filing Date
- 2023-08-08
- Publication Date
- 2025-08-26
AI Technical Summary
Current diagnostic methods for respiratory diseases, particularly in critical care settings, lack noninvasive and reliable molecular biomarkers for detecting respiratory tract infections (RTIs) in intubated patients, leading to delayed and inaccurate diagnoses, and there is a need for rapid, low-cost, and efficient methods to capture and analyze exhaled breath aerosols for disease detection.
A packed-bed column system connected to a ventilator exhalation tube captures cleaved proteoforms in exhaled aerosols, followed by mass spectrometry analysis to identify statistically significant proteoforms, using methods like microarray significance analysis and multiple logistic regression to predict RTIs with high accuracy.
The system provides rapid, accurate, and cost-effective diagnosis of RTIs by identifying characteristic proteoforms in exhaled breath, enhancing diagnostic specificity and reducing the need for invasive sampling.
Smart Images

Figure 2025528164000001_ABST
Abstract
Description
Related Applications
[0001] This patent application is an international application of U.S. Application No. 17 / 886,443, entitled "Capturing Cleaved Proteoforms in Exhaled Breath for Diagnosis and Treatment of Disease," filed on August 11, 2022, which is a continuation-in-part of U.S. Application No. 17 / 827,708, entitled "Capturing Cleaved Proteoforms in Exhaled Breath for Diagnosis and Treatment of Disease," filed on May 29, 2022, which is a continuation-in-part of International Application No. PCT / US2022 / 022964, filed on March 31, 2022, which is related to U.S. Provisional Application No. 17 / 886,443, which is entitled "Capturing Cleaved Proteoforms in Exhaled Breath for Diagnosis and Treatment of Disease," filed on March 31, 2021. This application is related to and claims priority from U.S. Provisional Application No. 63 / 169,130, entitled "Diagnosis of Respiratory Disease by Capturing Aerosolized Biomaterial Particles Using a Packed Bed System and Method," filed September 28, 2021; U.S. Provisional Application No. 63 / 249,357, entitled "Diagnosis of Respiratory Disease by Capturing Aerosolized Biomaterial Particles Using a Packed Bed System and Method," filed March 30, 2022; and U.S. Provisional Application No. 63 / 325,435, entitled "Diagnosis of Respiratory Disease by Capturing Aerosolized Biomaterial Particles Using a Packed Bed System and Method," filed March 30, 2022, the entire disclosures of which are incorporated herein by reference. FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] none. [Technical Field]
[0003] The present disclosure relates to methods and devices that use packed-bed columns to capture and analyze aerosolized organic biomaterials, such as viral and bacterial particles and associated truncated proteoforms, in exhaled breath, enabling rapid and low-cost detection of several diseases, including respiratory diseases such as COVID-19. More specifically, but not by way of limitation, the present disclosure relates to methods and devices that use mass spectrometry to analyze truncated proteoforms and non-volatile organic particles in exhaled breath for disease detection. [Background technology]
[0004] Exhaled aerosols contain nonvolatile organic biomarkers produced by human biological processes such as metabolism, immunity, and inflammation, and the composition of these compounds and proteoforms can be recognized as indicators of human health. Detection of these protein biomarkers and their cleaved proteoforms using exhaled breath analysis could enable monitoring, screening, diagnosis, and differentiation of healthy individuals from those with health problems such as obesity, diabetes, liver cancer, and lung cancer. Capturing these biomarkers from exhaled breath and subsequent analysis could reveal health risk factors and aid in the diagnosis, treatment, and mitigation of the spread of disease.
[0005] Although research has shown that respiratory diseases can be detected from exhaled aerosols and exhaled breath condensate, modern clinical testing for infections or diseases such as COVID-19, tuberculosis, influenza, and pneumonia continues to use sputum, blood, or nasal swabs. Coronavirus disease (COVID-19) is caused by the newly emerged coronavirus SARS-CoV-2. This novel coronavirus is a respiratory virus that spreads primarily through droplets produced when an infected person coughs or sneezes, or through droplets in saliva or nasal secretions. This novel coronavirus is highly contagious and has caused a pandemic. Furthermore, tuberculosis (TB) has surpassed HIV / AIDS as a global cause of death, killing more than 4,000 people per day (Patterson, B. et al., 2018). In areas with a high HIV prevalence, genotyping studies of Mycobacterium tuberculosis (Mtb) have found that the majority (54%) of TB cases are due to recent infection rather than reactivation. The physical processes of tuberculosis transmission remain poorly understood, and the application of new technologies to elucidate key events in the generation, release, and inhalation of infectious aerosols has been slow. Interrupting transmission would have a rapid and measurable impact on tuberculosis incidence. Rapid disease detection tools are needed to mitigate respiratory disease transmission.
[0006] The time associated with diagnostic assays is a critical parameter for field or "point-of-care" testing. Active case finding (ACF), by definition, occurs outside the healthcare system and is therefore an example of a field diagnostic assay. According to the World Health Organization, ACF is "the systematic identification of individuals suspected of having active tuberculosis using rapidly applicable tests, examinations, or other procedures." In the United States, point-of-care tests should provide a response, preferably within 20 minutes. Using the GeneXpert assay (Cepheid, Inc., Sunnyvale, California), a diagnosis can be provided in approximately one hour. The GeneXpert genetic assay is based on polymerase chain reaction (PCR) and can be used to analyze samples for the diagnosis of respiratory illnesses. This assay has not yet been widely adopted because it is expensive to implement on a "cost-per-test" basis. Due to its high cost, in developing countries, it is not used to screen apparently healthy (asymptomatic) patients for possible tuberculosis infection, but rather to confirm a diagnosis that is strongly suspected based on other tests or factors. The goal of ACF is to get infected individuals into treatment earlier, reducing the mean duration of infection and the spread of the disease. By the time an individual with tuberculosis seeks help at a clinic, they may have infected approximately 10 to 115 other people. ACF can help reduce or prevent significant tuberculosis transmission. Diagnostic systems and methods, such as sputum analysis and blood analysis, are not automated, operate autonomously, or are not rapid. Many systems and methods are expensive analytical methods that use expendable reagents per analysis, making them generally unsuitable for active case detection, especially in developing and less developed countries.
[0007] There is growing interest in novel diagnostic tools using exhaled breath for diseases, including respiratory disorders. Exhaled breath contains aerosols ("EBA") and vapors that can be collected noninvasively and characterized to elucidate physiological and pathological processes in the lungs (see Hunt, 2002). EBA analysis appears to be an attractive diagnostic tool for tuberculosis detection, due to its rapid analysis, portability, and low cost due to the lack of expensive assays and consumables. To capture exhaled breath for analysis, exhaled air is passed through a condenser, producing a liquid deposit called exhaled breath condensate (EBC). Although EBC is primarily composed of water vapor, it contains dissolved nonvolatile compounds such as cytokines, lipids, surfactants, ions, oxidation products, adenosine, histamine, acetylcholine, and serotonin. Additionally, EBC captures potentially volatile water-soluble compounds such as ammonia, hydrogen peroxide, ethanol, and other volatile organic compounds. EBC possesses an easily measurable pH. EBCs contain aerosolized airway lining fluid and volatile compounds that noninvasively indicate ongoing biochemical and inflammatory activity in the lungs. Interest in EBCs rapidly increased with the recognition that they possess measurable characteristics that can be used to distinguish infected from healthy individuals in pulmonary diseases. These analyses have provided evidence of airway and lung redox deviations, acid-base status, and the degree and type of inflammation in acute and chronic asthma, chronic obstructive pulmonary disease, adult respiratory distress syndrome, occupational diseases, and cystic fibrosis. Due to the uncertain and variable nature of dilution, EBCs may not accurately assess the concentrations of individual solutes within the native airway lining fluid. However, useful information can be provided when concentrations differ significantly between healthy and diseased individuals or when based on the ratio of solutes present in the sample.
[0008] Patterson et al. (2018) isolated and accumulated respiratory aerosols from a single patient using a respiratory aerosol sampling chamber (RASC), a novel device designed to optimize patient-derived exhaled aerosol sampling. Environmental sampling detected Mtb present after a period of aging in the air within the chamber. Thirty-five patients with newly diagnosed TB who had GeneXpert sputum-positive samples were confined and monitored in a RASC chamber with a volume of approximately 1.4 m3 for one hour. The GeneXpert TB PCR assay accepted sputum samples and provided positive or negative results in approximately one hour. The chamber incorporated aerodynamic particle size detection, viable and nonviable bacterial sampling, real-time CO2 monitoring, and cough recording. Microbial culture and droplet digital polymerase chain reaction (ddPCR) were used to detect Mtb in each bioaerosol collection device. Mtb was detected in 77% of aerosol samples, with 42% of samples positive by Mtb culture and 92% positive by ddPCR. A correlation was observed between cough rate and culturable bioaerosols. Mtb was detected at all viability cascade impactor stages, with peaks in aerosol sizes between 2.0 and 3.5 microns. This suggests a median of 0.09 CFU per liter of exhaled breath for aerosol culture positivity and an estimated median exhaled particulate bioaerosol concentration of 4.5 x 107 CFU / ml. Mtb was detected in bioaerosols exhaled by the majority of untreated TB patients using the RASC chamber. Molecular detection was found to be more sensitive than Mtb culture on solid media. Breath analysis tools have not been commercialized for ACF due to a lack of methods and devices to efficiently collect and concentrate trace amounts of analytes present in exhaled breath. Furthermore, there is no standard or methodology for assessing the volume of exhaled breath sufficient for a specific diagnosis.
[0009] The lack of noninvasive methods and reliable molecular biomarkers poses a significant barrier to the diagnosis of respiratory tract infections (RTIs) in critical care settings, particularly in mechanically ventilated patients. Current diagnostic methods rely on nonspecific clinical observations, such as tracheal secretions, chest radiographs, temperature, white blood cell count, oxygenation, and microbiological testing. Scoring systems, such as the Clinical Pulmonary Infection Score (CPIS), have been developed based on these clinical manifestations. Although clinical records and scoring systems can be used to determine antibiotic treatment, they generally lack sensitivity and specificity for RTI diagnosis, making it difficult for clinicians to make rational clinical decisions. Quantitative microbial cultures of specimens collected from the lower respiratory tract, such as noninvasive endotracheal aspirates (ETAs), have been used to diagnose RTIs, but they cannot determine whether the identified bacteria are due to general airway colonization or a separate infection. Bronchoalveolar lavage (BAL) is used as a high-quality specimen collection technique from the lower respiratory tract for etiologic diagnosis in intubated patients. However, this method is invasive and cannot be routinely performed in hospital ICUs. Due to these limitations, more than 50% of patients in intensive care units (ICUs) are treated without an appropriate diagnosis. Therefore, current diagnostic methods, pathogen identification, and management of RTIs in intubated patients are limited by the difficulty of collecting samples from the infection site and the lack of accurate diagnostic molecular biomarkers. There is an urgent need to develop noninvasive methods for sampling infection sites and discovering accurate molecular biomarkers for RTI diagnosis.
[0010] Noninvasive sampling methods allow for repeated sampling without risk to critically ill patients, allowing for disease progression monitoring. Direct sampling from the lower respiratory tract provides specimens that more accurately represent the site of infection, improving diagnostic specificity. Noninvasive sampling methods also facilitate patient participation in clinical trials, which are beneficial for treatment and diagnostic research. Human respiratory and exhaled aerosols have the potential to be used as noninvasive sources for clinical use. Organic molecules contained in human respiratory and exhaled aerosols could potentially be used to develop noninvasive methods for detecting lung disease exacerbations and infections. Organic molecules contained in human breath include two main types: volatile organic compounds (VOCs) and nonvolatile organic compounds (NOCs). VOCs are gas molecules emitted from nonbiological sources, such as diet, plants, and household cleaners, and therefore lack specificity for use as biomarkers. NOCs, on the other hand, are large molecules derived exclusively from living organisms, either humans or pathogens, making them more suitable as surrogate biomarkers. Noninvasive sampling methods targeting NOCs are being developed for use in clinical settings. McNeil et al. reported the use of an in-line heat and moisture exchanger (HME) filter to collect proteins from patients with acute respiratory distress syndrome (ARDS). HME filters are standard components attached to mechanical ventilators where exhaust is present. They reported that proteins can be trapped on HME filters as exhaled condensate released from the lower airways. To this end, they collected undiluted pulmonary edema fluid (EF) samples and compared the protein profiles obtained from the EF samples with those from HME fluid samples. The results showed similar protein profiles between the two types of samples, suggesting that HME could be a noninvasive alternative to EF for distal airspace sampling in ARDS patients.
[0011] HME filters have limitations. They contain a hygroscopic, sponge-like material. Protein capture is presumed to occur via condensation on the sponge-type material. During condensation, Reifart et al. (2021) reported that submicron particles, such as the SARS-CoV-2 virus, are not efficiently collected on filters, primarily because the particles in human exhaled breath are too small (<1 μm). Because particles in human respiratory and exhaled aerosols are primarily composed of submicron particles, capturing these particles using the disclosed exemplary devices and methods overcomes the limitations of HME filters by collecting exhaled aerosol and exhaled condensate in relatively concentrated samples with high flow rates and high efficiency. Furthermore, the disclosed exemplary devices and methods provide sample normalization by enabling recording of individual CO2 levels in exhaled breath.
[0012] Additionally, incorporating aerosol size sorting can enhance the signal-to-noise ratio for certain analytes prior to analyte collection. The enriched sample is then analyzed by several methods, preferably using methods that are sensitive, rapid, and highly specific for the analyte of interest. More preferably, the analysis is rapid and near real-time. Mass spectrometry, real-time PCR, and immunoassays have the greatest potential for sensitivity, specificity, and near real-time analysis. There is a need for sample collection methods that can be coupled with rapid diagnostic tools such as mass spectrometry ("MS"), which are faster and more reliable than sputum analysis and less invasive than blood analysis, to provide diagnostic assays that are rapid, sensitive, specific, and preferably have low cost per test. Such systems can be used for active case finding ("ACF") of respiratory diseases, as well as to monitor the status of patients receiving mechanical ventilators to support their breathing in hospital intensive care units. To be effective, sample collection and diagnostic systems must be fast and inexpensive "per diagnosis." A low cost per test is essential for screening a large number of individuals to proactively prevent disease transmission and identify the few who are actually infected. Low-cost devices and methods are also needed for point-of-care diagnostics of influenza and other pathogenic viruses, as patients who may have a "common cold" may also be infected with rhinovirus. Respiratory infections are sometimes caused by bacterial or fungal microorganisms and can be treated with antibiotics. In other cases, microorganisms may be resistant to antibiotics, and diagnostic methods that can identify microbial resistance to antibiotics are desirable. Rapid EBA methods are needed to distinguish viral from bacterial infections in the respiratory tract while minimizing the possibility of false negatives due to insufficient sample volume. Mass spectrometry, genomics methods including PCR, and immunoassays have the potential to offer the highest sensitivity and specificity. Mass spectrometry, particularly MALDI time-of-flight mass spectrometry (MALDI-TOFMS), has proven to be sensitive, specific, and near-real-time, making it the preferred diagnostic tool for analyzing EBA and EBC samples. Summary of the Invention
[0013] An exemplary method for predicting respiratory tract infections (RTIs) in intubated patients breathing with the assistance of a ventilator is disclosed. The method includes culturing one or more of a sputum sample, an endotracheal tube sample ("ET"), or a bronchoalveolar lavage fluid ("BAL") to diagnose the presence or absence of an RTI and obtain baseline data for each patient in a clinical study, regardless of whether or not the patient has an RTI. The method may continue by selectively capturing cleaved proteoforms in the exhaled aerosols generated by each patient using a packed-bed column removably connected to the exhalation tube of the ventilator; extracting the cleaved proteoforms from the packed-bed column into one or more collected liquid samples corresponding to each patient; and analyzing the one or more collected liquid samples containing the cleaved proteoforms using mass spectrometry to obtain raw mass spectra. The method may continue with predicting the presence of an RTI using one or more of the following steps: identifying a statistically significant subset of truncated proteoforms characteristic of an RTI; calculating a composite score representing the statistically significant subset of truncated proteoforms; or calculating an area under the curve (AUC) of a receiver operating characteristic (ROC) curve representing the statistically significant subset. Identifying the statistically significant subset of truncated proteoforms may include identifying statistically significant classes of truncated proteoforms characteristic of an RTI in the mass spectra using a mass spectral feature selection method (including one or more of microarray significance analysis (SAM) ranking or t-test) with reference to baseline data; and refining the statistically significant subset of classes of truncated proteoforms using multiple logistic regression analysis of variables including one or more of age, sex, race, ethnicity, primary diagnosis, medication, sample collection time, microorganism identification information, white blood cell count, body temperature, fraction of inspired oxygen (FiO2) content, lung radiograph, or truncated proteoforms within the class.
[0014] In an exemplary method, identifying statistically significant classes of truncated proteoforms using a t-test can include applying a two-tailed unpaired t-test to the truncated proteoforms and adjusting the p-value using the Benjamini-Hochberg method to apply a false discovery rate ("FDR") of 0.05. The refinement step can include selecting truncated proteoforms with p-values less than 0.05 obtained from the multiple logistic regression analysis to generate a statistically significant subset of truncated proteoforms. Predicting the presence of RTI by calculating a composite score representing a statistically significant subset of truncated proteoforms can include using a reference data sample containing a statistically significant subset of truncated proteoforms; determining a reference threshold mass spectral intensity value for each truncated proteoform in the subset as a value equal to the normalized mass spectral intensity value (log10) associated with the intersection of the specificity curve and the sensitivity curve in the ROC for each proteoform; assigning an index score of 1 to a truncated proteoform in the subset if the measured mass spectral intensity value (log10) of the truncated proteoform is equal to or greater than the reference threshold intensity value, and assigning an index score of 0 to a proteoform if the measured mass spectral intensity value of the truncated proteoform is less than the reference threshold intensity value; determining a cutoff classification value representing the minimum number of statistically significant truncated proteoforms in the subset; adding the index scores assigned to each statistically significant truncated proteoform in the subset to calculate a composite score representing a statistically significant subset of truncated proteoforms for each sample collected; and predicting the presence of RTI if the composite score is equal to or greater than the cutoff classification value.
[0015] In an exemplary method, determining the cutoff classification value may include generating a confusion matrix for each classification value (n, (n-1), (n-2), ..., 0) (where n is the total number of statistically significant proteoforms in the subset using the index score (0 or 1) of each proteoform as a predictive index and using the baseline data as an actual index of RTI (0 or 1)), calculating the RTI prediction accuracy for each classification value using the confusion matrix, defined as the ratio of the sum of true positive results and true negative results to the total number of fluid samples collected, and determining the cutoff classification value as the classification value that includes the number of truncated proteoforms required to yield an RTI prediction accuracy of at least about 90%. The exemplary method may further include determining whether a composite score is statistically significant in distinguishing between RTI and non-RTI patients when a p-value of the composite score obtained from a multiple logistic regression analysis of variables including one or more of age, sex, race, ethnicity, primary diagnosis, medication, sample collection time, microbial identity, white blood cell count, body temperature, fraction of inspired oxygen (FiO2) content, lung radiograph, individual scores of truncated proteoforms within the subset, or the composite score is less than 0.001. The exemplary method may further include predicting the presence of an RTI by calculating the area under the curve (AUC) of a receiver operating characteristic (ROC) curve representing all proteoforms within the statistically significant subset of truncated proteoforms. This step may include constructing a ROC representing all proteoforms in the statistically significant subset (wherein the index score for each proteoform is used as a predictive indicator of RTI and the baseline data is used as an actual indicator of RTI to calculate specificity and sensitivity values for the ROC), determining the area under the curve (AUC) using the ROC representing all proteoforms in the statistically significant subset, and predicting the presence of RTI if the AUC value is greater than at least about 95%.
[0016] An exemplary method for diagnosing respiratory tract infections ("RTIs") in intubated patients by capturing cleaved proteoforms in exhaled aerosols is disclosed, the diagnostic method comprising the steps of selectively capturing cleaved proteoforms in exhaled aerosols produced by each patient using a packed bed column removably connected to the exhalation tube of a ventilator, extracting the cleaved proteoforms into one or more collected liquid samples corresponding to each patient, analyzing the collected samples corresponding to each patient containing the cleaved proteoforms using mass spectrometry to obtain raw mass spectra, calculating a composite score of statistically significant proteoforms in the samples, where the statistically significant proteoforms are provided by reference data as described above, and diagnosing the presence of RTI if the composite score is equal to or greater than a composite score in the reference data that predicts RTI with at least greater than 90% accuracy. Calculating a composite score for statistically significant proteoforms in a sample may include determining the normalized mass spectral intensity value (log10) of each statistically significant truncated proteoform, assigning an index score of 1 to the statistically significant truncated proteoform if the normalized intensity value of the truncated proteoform is equal to or greater than its reference threshold intensity value, and assigning an index score of 0 to the statistically significant truncated proteoform if the normalized intensity value of the proteoform is less than its reference threshold intensity value, and adding the index scores to calculate a composite score representing a statistically significant subset of truncated proteoforms in the sample.
[0017] An exemplary packed-bed column may include one or more of resin beads having C18 functional groups on their surface, cellulose beads having sulfate ester functional groups on their surface, or a mixture thereof. The resin beads and cellulose beads may have a nominal diameter of at least about 20 μm. The resin beads and cellulose beads may have a nominal diameter between about 40 μm and about 150 μm. Extracting the cleaved proteoforms may include flushing the packed-bed column with at least one solvent and collecting the solvent containing the cleaved proteoforms from the packed bed. The at least one solvent may include one or more of acetonitrile, methanol, trifluoroacetic acid (TFA), or isopropanol (IPA), with the remainder being water. The one or more solvents may include about 50% to about 70% acetonitrile, about 50% to about 70% isopropanol, or about 0.05% TFA by volume in water. A statistically significant subset of the class of truncated proteoforms may include one or more of CO6A3 (amino acids 2781-2792), CYTA (2-17), DEN2B (628-637), IRAK4 (121-130), MMP9 (673-691), or PHTF2 (271-285).
[0018] An exemplary breath collection system for capturing cleaved proteoforms in exhaled aerosols for disease diagnosis and treatment is disclosed. The exemplary system may include one or more sample capture elements including a packed-bed column that selectively captures aerosolized cleaved proteoforms in patient-generated exhaled breath, and a subsystem configured to be fluidly and electrically connected to the sample capture element using a quick connect / disconnect coupling. The subsystem may include one or more of a pump that draws the exhaled aerosol into the sample capture element, a power source, or a controller that controls the operation of the sample capture element. The one or more sample capture elements may be removably connected to an exhalation tube of a ventilator used to assist the breathing of intubated patients. The controller may be configured to detect proper mechanical and electrical contact between the sample capture element and the subsystem and alert a user via one or more graphical user interfaces or audible alarms disposed in the subsystem. The subsystem may further include one or more CO2 sensors or particle counters disposed between the sample capture element and the pump. The subsystem may further include a trap disposed between the one or more sample capture elements and the pump and configured to capture exhaled breath condensate (EBC) containing one or more water vapor, volatile organic components, or nonvolatile organic components passing through the packed bed. The packed-bed column may include solid particles including one or more resins, cellulose, silica, agarose, or hydrated Fe3O4 nanoparticles. The packed-bed column may include resin beads having C18 functional groups on their surfaces, cellulose beads having sulfate ester functional groups on their surfaces, or a mixture thereof. The resin beads and cellulose beads may have a nominal diameter of at least about 20 μm. The resin beads and cellulose beads may have a nominal diameter of about 40 μm to about 150 μm. The resin beads may be packed between two porous polymer frit disks. The nominal flow rate drawn through the bed using the pump may range from about 200 mL / min to about 3 L / min.
[0019] An exemplary system for capturing cleaved proteoforms in exhaled breath to diagnose and treat disease is disclosed, which includes the exhaled breath collection system described above, a sample extraction system for extracting the captured cleaved proteoforms characteristic of the disease from the packed bed column into one or more liquid samples, and an analytical device for analyzing the cleaved proteoforms in the one or more liquid samples. The extraction system can include means for flushing the packed bed column with at least one solvent and collecting the solvent containing the cleaved proteoforms from the packed bed. The analytical device can include one or more of PCR, ELISA, rt-PCR, mass spectrometry (MS), MALDI-MS, ESI-MS, or MALDI-TOF MS, and LC-MS / MS.
[0020] An exemplary method for predicting the presence of disease by capturing cleaved proteoforms in exhaled aerosol is disclosed, which includes the steps of diagnosing the presence or absence of disease and obtaining baseline data for each patient in a clinical laboratory study by culturing one or more of a sputum sample, an endotracheal tube sample ("ET"), or a bronchoalveolar lavage fluid ("BAL") sample. The method may continue with the steps of selectively capturing cleaved proteoforms in the exhaled aerosol produced by each patient using a packed-bed column, extracting the cleaved proteoforms from the packed-bed column into one or more collected liquid samples corresponding to each patient, analyzing the one or more collected liquid samples containing the cleaved proteoforms using mass spectrometry to obtain raw mass spectra, and identifying a statistically significant subset of cleaved proteoforms characteristic of disease. The method may continue with predicting the presence of disease using one or more of the following steps: calculating a composite score representing a statistically significant subset of truncated proteoforms; or calculating an area under the curve ("AUC") of a receiver operating characteristic curve ("ROC") representing a statistically significant subset. Identifying the statistically significant subset of truncated proteoforms may include identifying statistically significant classes of truncated proteoforms characteristic of disease within the mass spectra using mass spectral feature selection methods, including one or more of SAM (Significance Analysis of Microarrays) ranking or t-tests, with reference to baseline data; and refining the statistically significant subset of classes of truncated proteoforms using multiple logistic regression analysis of variables, including one or more of age, sex, race, ethnicity, primary diagnosis, medication, sample collection time, microbial identity, white blood cell count, body temperature, fraction of inspired oxygen (FiO2) content, lung radiograph, or truncated proteoforms within the class.
[0021] Other features and advantages of the present disclosure will be set forth in part in the following description and the accompanying drawings, in which preferred aspects of the present disclosure are described and shown, and in part will become apparent to those skilled in the art through examination of the following detailed description, taken in conjunction with the accompanying drawings, and will be learned through the practice of the present disclosure. The advantages of the present disclosure may be realized and attained by means of the instrumentalities and combinations particularly pointed out in the appended claims. [Brief explanation of the drawings]
[0022] The foregoing aspects and many of the attendant advantages of the present disclosure will be more readily appreciated as the same become better understood by reference to the following detailed description, when taken in conjunction with the accompanying drawings. [Figure 1A] FIG. 1A is a schematic diagram of an exemplary exhaled aerosol collection system for use with a ventilator connected to a patient diagnosed with a respiratory tract infection (RTI) in an intensive care unit. [Figure 1B] FIG. 1B is a schematic diagram of an exemplary subsystem configured to operate an exhaled aerosol sample capture system connected to a ventilator. [Figure 2] FIG. 2 is a schematic diagram of an exemplary respiratory disease diagnostic system including an exemplary breath sample collection system. [Figure 3A] Figure 3A shows boxplots for distinguishing RTI and non-RTI patients using 263 truncated proteoform classes identified using mass spectrometry of exhaled aerosol. [Figure 3B] Figure 3B shows the distribution (volcano plot) of feature ranking scores and fold changes of the six statistically significant truncated proteoforms for distinguishing RTI and non-RTI patients based on the ion intensities of the six truncated proteoforms. [Figure 3C] Figure 3C shows boxplots for distinguishing between RTI and non-RTI patients using select classes of six truncated proteoforms identified using mass spectrometry of exhaled aerosol. [Figure 3D]Figure 3D is an estimate of the reference threshold mass spectral intensity values (log10) for each of the three statistically significant truncated proteoforms from the respective ROC curves (a-c). [Figure 4] FIG. 4 is a schematic diagram of an exemplary method for predicting RTI from the mass spectrum of an exhaled aerosol sample collected from a patient using an exemplary exhaled aerosol sample capture element and collection system. [Figure 5A] Figure 5A shows a plot depicting the relationship between the composite score estimated for the three statistically significant truncated proteoforms and the probability of distinguishing between RTI and non-RTI (RTI prediction accuracy). [Figure 5B] Figure 5B shows the ROC curves showing the AUC values for each of the three truncated proteoforms. [Figure 5C] Figure 5C shows the ROC curve with AUC values for the general linear model using a selected subset (three proteoforms) of the six truncated proteoform classes.
[0023] All reference numbers, identifiers, and callouts in the figures are incorporated herein by this reference as if fully set forth herein. Failure to number elements in the figures is not intended as a waiver of any rights. Non-numbered references may also be identified by the alphabetic letter of the figure or appendix.
[0024] The following detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, specific embodiments in which the disclosed systems and methods may be practiced. These embodiments, which may be understood as "examples" or "options," are described in sufficient detail to enable one skilled in the art to practice the invention. The embodiments may be combined, other embodiments may be utilized, or structural or logical changes may be made without departing from the scope of the invention. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the invention is defined by the appended claims and their legal equivalents.
[0025] In this disclosure, aerosol generally refers to a suspension of particles dispersed in air or gas. "Autonomous" diagnostic systems and methods mean that they generate diagnostic test results "without or minimal intervention by a medical professional." The U.S. FDA classifies medical devices based on the risk associated with the device and by assessing the amount of regulation that reasonably ensures the device's safety and effectiveness. Devices are classified into one of three regulatory classes: Class I, Class II, or Class III. Class I includes devices with the lowest risk, while Class III includes devices with the highest risk. All classes of devices are subject to general regulation, a fundamental requirement of the Food, Drug, and Cosmetic (FD&C) Act that applies to all medical devices. In vitro diagnostic products are reagents, instruments, and systems intended for use in the diagnosis of disease or other conditions, including the determination of health status, in order to cure, mitigate, treat, or prevent disease or its sequelae. Such products are intended for use in the collection, preparation, and testing of specimens taken from the human body. The exemplary devices disclosed herein operate autonomously and can generate reliable results, thereby potentially being regulated as Class I devices. In some parts of the world with a high burden of TB infection, access to medically trained personnel is very limited. Autonomous diagnostic systems are preferred over non-autonomous diagnostic systems.
[0026] In this disclosure, the singular (equivalent to the English terms "a" or "an") is used to include one or more, and the term "or" is used to refer to a non-exclusive "or" unless otherwise specified. Furthermore, it is to be understood that expressions or terms used herein and not otherwise defined are for descriptive purposes only and not for limiting purposes. Unless otherwise specified in this disclosure, for interpreting the scope of the term "about," the margin of error associated with disclosed values (dimensions, operating conditions, etc.) is ±10% of the value set forth in this disclosure. The margin of error associated with values disclosed as percentages is ±1% of the stated percentage. The word "substantially" used before certain words includes the meanings "a significant portion of the specified range" and "most but not all of what is specified." Unless otherwise specified, concentrations of chemicals, solvents, etc. disclosed as percentages refer to volume %. Detailed Description
[0027] Exhaled aerosol particles contain various nonvolatile organic biomolecules, such as metabolites, lipids, and proteins. Exhaled aerosol particles may contain one or more of the following: microorganisms, viruses, metabolic biomarkers, lipid biomarkers, or proteomic biomarkers (e.g., truncated proteoforms characteristic of respiratory and other diseases). Furthermore, these nonvolatile molecules have a wide particle size distribution, ranging from submicron to approximately 10 microns. There is a need for breath collection and disease diagnosis systems and methods that can efficiently capture various types of nonvolatile molecules of various particle sizes from exhaled breath. Certain aspects of the present invention are described in considerable detail below for the purpose of explaining the configuration, principles, and operation of the disclosed methods and systems. However, various modifications can be made, and the scope of the present invention is not limited to the exemplary aspects described. An example of a noninvasive method for distinguishing between RTI and non-RTI patients by capturing truncated proteoforms in the exhaled aerosol of intubated patients is disclosed.
[0028] An exemplary system 1300 (FIG. 1A) and method for capturing exhaled aerosols is disclosed, in which an exemplary sample capture element 1301, including a packed-bed column, is placed in fluid communication with a ventilator 1305. The ventilator 1305 is a life-support device used in intensive care units for patients unable to breathe on their own. For example, patients with severe COVID-19 symptoms may require ventilator assistance to breathe. A tube 1306 is inserted through the patient's mouth or nose directly into the trachea. The ventilator pumps air into the lungs through this tube, forcing the patient to inhale. The ventilator typically pumps air for one second, pauses for approximately three seconds, and then repeats the cycle to allow the patient to exhale through the same tube. The inlet end 1302 of the capture element 1301 is removably connected, preferably directly to the exhalation tube of the ventilator, to minimize particle loss. The outlet end 1303 may be removably connected to a pump 1308 of subsystem 1304 (details of which are shown in subsystem 1313), which may use tubing to draw exhaled air through the packed bed column of element 1301 at a flow rate of approximately 200 ml / min to approximately 2.5 L / min. System 1300 may include a trap disposed between end 1303 and subsystem 1313 to collect condensate. The trap may be cooled to a temperature below ambient temperature. An optional HEPA filter and needle valve or flow meter may be implemented between the trap and the pump. CO2 in the exhaled air passes through the packed bed column. A CO2 sensor may be disposed between the outlet end 1303 and the trap to determine if the exhaled air sample volume is adequate. Monitoring CO2 allows for an estimate of the exhaled air volume. A particle counter may also be placed upstream of the capture element 1301 and between the outlet end 1303 and the trap to detect the size and number of particles exiting the packed bed column, and the particle counter may be used to detect saturation of the packed bed and permeation of non-volatile organic molecules from the column bed.
[0029] The sample capture element 1301 can include a packed-bed column that selectively captures nonvolatile respiratory aerosol particles. The capture element 1301 can be placed in fluid communication with the system 1313 (FIG. 1B) via a port 1314. The port 1314 can include a quick connect / disconnect coupling. A portion of the exhaled air drawn through the capture element 1301 using the pump 1308 can be delivered to a reservoir 1312 that is fluidly connected to a CO2 sensor 1311. The reservoir 1312 can be a tightly sealed container and is used to prevent air leakage from the CO2 sensor. The system 1313 can include a user interface and an on / off switch for starting and stopping exhaled air sampling using the element 1301. Additionally, components such as a flow controller and a flow restrictor 1309 can also be packaged in the portable subsystem 1313. The subsystem 1304 can include a diaphragm pump, such as a mini diaphragm pump 1308. The portable system 1313 is 11 inches x 7.5 inches x 5.5 inches (length x depth x height) and can include noise-reducing materials such as foam padding to reduce the noise level generated by the pump to less than 45 dB. The system 1313 can be placed remotely from the sample capture element, for example, outside of a hospital intensive care unit.
[0030] An exemplary packed-bed column of capture element 1301 can contain Hamilton PRP-C18 resin beads supplied by Sigma-Aldrich and other vendors. The bed can be held between two porous filter plates, such as fritted disks. For example, a polyethylene disk with an average pore size greater than 35 μm can be placed upstream of the bed, and a polyethylene disk with an average pore size of 10 μm (Boca Scientific, Dedham, Massachusetts) can be placed downstream of the bed. The 35 μm fritted disk allows for a faster airflow rate, while the smaller 10 μm fritted disk effectively captures all of the C18 resin. In exemplary element 1301, the packed bed can contain only about 25 mg of C18 resin beads with a nominal diameter of about 12 μm to about 20 μm. Nonvolatile organic components in the exhaled breath removably interact with the C18 functional groups on the beads and are captured. Water, volatiles, and other hydrophilic molecules can pass through the bed and be captured in a glass trap.
[0031] In addition to C18 functional groups, other functional groups exhibiting affinity for nonvolatile molecules can be used as adsorbents in columns immobilized on solid phase beads, such as resin beads. Solid phase beads can be made from polymers and particles, such as resin, cellulose, silica, agarose, and hydrated Fe3O4 nanoparticles. The adsorbent material can include other functional groups deposited on the solid phase beads, including, but not limited to, octadecyl, octyl, ethyl, cyclohexyl, phenyl, cyanopropyl, aminopropyl, 2,3-dihydroxypropoxypropyl, trimethylaminopropyl, carboxypropyl, benzenesulfonic acid, and propylsulfonic acid. Functional groups can also include one or more of an ion exchange phase, a polymer phase, an antibody, a glycan, a lipid, DNA, or RNA. To capture aerosolized virus particles, an exemplary sample capture element 1301 can include sulfate-immobilized cellulose beads. Alternatively, the sample capture element 1301 can include a packed bed of C18 beads and sulfate-immobilized cellulose beads. Alternatively, the sample capture element 1301 may include a packed bed of a mixture of C18 beads and sulfate-immobilized cellulose beads. Exemplary sulfate beads may include Cellufine sulfate beads (JKC Corp., Japan). The particle diameter may be about 40 μm to about 130 μm. An exemplary sample capture element may include about 100 mg of sulfate-immobilized cellulose beads arranged as a packed bed column. An exemplary sample capture element may have an inner diameter of about 7 mm and a length of about 30 mm.
[0032] The nonvolatile organic molecule capture capacity of the C18 beads in element 1301 can be between about 0.05 mg (nonvolatile organics) / mg beads and about 0.5 mg / mg. The capacity of the C18-bonded resin beads in the column bed in an exemplary capture element can be about 0.1 mg / mg. That is, a column bed comprising 25 mg of C18 beads is expected to be characterized by a capture or adsorption capacity of about 2.5 mg of nonvolatile organic molecules. Pump 1308 can be a diaphragm pump. Data from the CO2 sensor is recorded on a nonvolatile memory card, such as an SD card commonly used in portable devices. A flow sensor can be installed to monitor the flow rate through the C18 packed bed column. Alternatively, a flow controller can be used to achieve a constant flow rate, e.g., 500 mL / min, through the packed bed column. To use the example capture element 1301 to sample exhaled aerosols from a ventilator installed in a hospital intensive care unit, the pump 1308 may be packaged into a portable system 1313 along with a CO2 sensor 1311, associated power supply 1307, system control components, and necessary fluidic components (tubing, quick connect / disconnect couplings, etc.) (FIG. 1B).
[0033] An exemplary diagnostic system 2000 (FIG. 2) is disclosed, which can include a breath sample collection system 2001 disposed in fluid communication with a sample extraction system 2002 and an analysis system 2003. The sample collection system 2001 can include the exemplary collection system 1301 described above. After a predetermined sample collection period, the sample capture element 1301 can be removed from the system 1300. The element 1301 can then be autoclaved at 110° C. for about 10 minutes to sterilize the element 1301 before extracting the captured aerosol particles. The captured non-volatile aerosol particles can be extracted by washing (or flushing) the column with about 200 μL to about 400 μL of a solvent including one or more of 70% acetonitrile (ACN), about 50% to about 70% methanol, or about 50% to about 70% isopropyl alcohol (IPA). For example, a 50% ACN flush may be used to elute metabolites and proteins in the first flush, followed by a 70% IPA flush to elute lipids from the packed-bed column. To preserve the captured bioaerosol particles, the organic solvent may be removed from the packed-bed column by lyophilization overnight, if desired. The organic solvent may be removed by incubating on a heating block at about 70°C for about 30 minutes. Finally, the bed may be washed with about 0.05% TFA (trifluoroacetic acid). A sample extraction system may be used to extract captured non-volatile organics from the packed-bed column of system 1300 and may be configured inline or offline within the system. If system 2002 is configured offline, at the end of breath sample collection, capture element 1301 is removed from system 1300 and eluted with organic solvent in extraction system 2002 to remove non-volatile organics from the packed-bed column. Examples of organic solvents for extracting trapped non-volatile organics (such as highly polar non-volatile organic molecules, proteins, etc.) from a packed bed column include, but are not limited to, about 50-70% acetonitrile in water.The extraction can be repeated using the same or a different solvent, including, but not limited to, 50-70% isopropanol in water to extract less polar lipid molecules from the packed bed. Other organic solvents include about 50% to about 70% methanol in water and about 50% methanol in about 50% chloroform. When system 2002 is configured in-line, one or more CO sensors or particle counters can be located upstream of extraction system 2002. System 2002 can include a solvent container, a pump for transferring the solvent from the solvent to the packed bed column, and a container for collecting the solvent containing non-volatile biomarkers in a separate container or cup. Alternatively, system 2002 can include an injector for injecting the solvent into the packed bed column and collecting the extract containing non-volatile organics and biomarkers in a suitable cup or container or other small-volume laboratory tube. The sample captured in the solvent can be further processed and analyzed in analysis system 2003.
[0034] Many diagnostic instruments can be adapted for use in the analytical system 2003, including, but not limited to, instruments that perform genome-based assays (e.g., PCR, rt-PCR, whole genome sequencing), biomarker recognition assays (e.g., ELISA), and spectral analyses such as mass spectrometry (MS). Of these diagnostic instruments, MS is preferred due to its analytical speed. Preferred MS techniques for biomarker identification are electrospray ionization (ESI) and matrix-assisted laser desorption / ionization (MALDI) time-of-flight MS (TOFMS). ESI can be coupled with a high-resolution mass spectrometer. MALDI-TOFMS instruments are small and lightweight, consume less than 100 watts, and can provide sample analysis in less than 15 minutes. MALDI-TOFMS is a suitable diagnostic instrument for point-of-care diagnostics suitable for ACF. The sample must be inserted into the vacuum chamber of the MS and dried before receiving a laser pulse from an ultraviolet laser. The interaction of the sample and laser generates large, information-rich biological ion clusters that are characteristic of the biological material. If the sample processing system 2004 provides a concentrated sample containing only trace amounts of water or organic solvents (such as acetonitrile, methanol, or isopropanol in a 50% to 70% aqueous solution), sample analysis using MS may take less than 5 minutes (including sample preparation) due to the short time required to evaporate the water from the sample.
[0035] MALDI-TOFMS has been used to detect Bacillus anthracis spores (multiple strains), Yersinia pestis, Francisella tularensis, Venezuelan equine encephalitis virus (VEE), Western equine encephalomyelitis virus (WEE), Eastern equine encephalitis virus (EEE), botulinum neurotoxins (BoNT), Staphylococcal enterotoxin (SEA), Staphylococcal enterotoxin B (SEB), ricin, abrin, Ebola Zaire strain, aflatoxin, saxitoxin, conotoxin, Enterobacterial phage T2 (T2), and HT-2. The presently disclosed exemplary systems and methods may be used to identify non-volatile biochemical threats such as HT2, cobra toxin, B. globizii spores, B. cereus spores, B. thuringiensis al. hakam spores, B. anthracis Sterne spores, Yersinia enterocolitica, E. coli, MS2 virus, T2 virus, adenovirus, and non-volatile drugs such as NGA (non-volatile), bradykinin, oxytocin, substance P, angiotensin, diazepam, cocaine, heroin, fentanyl, etc. Additionally, the exemplary systems and methods disclosed herein may be used to achieve accurate detection and identification of SARS-CoV-2 from human breath samples.
[0036] In matrix-assisted laser desorption / ionization (MALDI), target particles (analytes) are coated with a matrix chemical that preferentially absorbs light from a laser (often ultraviolet wavelengths). In the absence of a matrix, biological molecules decompose by pyrolysis when exposed to a laser beam in a mass spectrometer. The matrix chemical also transfers charge to the vaporized molecules, generating ions that are accelerated down a flight tube by an electric field. Microbiology and proteomics have become major application areas for mass spectrometry. Examples include bacterial identification, chemical structure discovery, and derivation of protein function. MALDI-MS has also been used for lipid profiling of algae. In MALDI-MS, a liquid containing an acid, such as trifluoroacetic acid (TFA), and a MALDI matrix chemical, such as α-cyano-4-hydroxycinnamic acid, dissolved in a solvent is added to the sample. Solvents include acetonitrile, water, ethanol, and acetone. TFA is typically added to reduce the effect of salt impurities on the sample's mass spectrum. Water allows for the dissolution of hydrophilic proteins, while acetonitrile allows for the dissolution of hydrophobic proteins. The MALDI matrix solution is spotted onto the sample on the MALDI plate to obtain a uniform, homogeneous layer of MALDI matrix material on the sample. The solvent evaporates, leaving only the recrystallized matrix, and the sample spreads throughout the matrix crystals. The acid partially decomposes the sample's cell membrane, making the proteins available for ionization and analysis in MS. Other MALDI matrix materials include 3,5-dimethoxy-4-hydroxycinnamic acid (sinapic acid), α-cyano-4-hydroxycinnamic acid (α-cyano or α-matrix), and 2,5-dihydroxybenzoic acid (DHB), as described in U.S. Patent No. 8,409,870.
[0037] Analytical methods for metabolites, proteins, and lipids can include silver staining for protein profiling, protein assays for protein content, bottom-up proteomics and LC-MS / MS for metabolite and lipid omics, and MALDI-TOF mass spectrometry for molecular profiling. In an exemplary study, exhaled aerosols from a patient infected with pneumonia were collected using a capture element 1301 connected to a ventilator. Subsequent analysis revealed that protein content measured using a protein assay and molecular profiling measured using MALDI-TOF MS were good indicators of the patient's pneumonia infection, as revealed by a Pearson correlation heat map including the variables of total exhaled air volume collected, exhaled CO2 content, protein content, MALDI-TOF total ion intensity, and MALDI-TOF MS single peak (4820 m / z) intensity.
[0038] The analytical system 2003 may include a sample processing system 2004 and one or more diagnostic devices 2005. The sample processing system 2004 may include the elements necessary to perform one or more of the following steps:
[0039] (a) Placing the sample into at least one of a cup, vial, or sample plate. For example, the Series 110A Spot Sampler (aerosol device) uses a 32-well plate with circular wells (75 μL well volume) or teardrop-shaped wells (120 μL well volume), which is heated to evaporate the solvent and excess fluid / liquid within the sample, concentrating the sample.
[0040] (b) placing the sample in a cup and exposing it to a vacuum source or a freeze-drying device to evaporate the solvent and concentrate the sample; and
[0041] (c) High temperature digestion of proteins and virus particles.
[0042] The sample may be centrifuged to remove chemical contaminant particles.
[0043] Virus (e.g., SARS-CoV-2) detection relies on the challenging task of detecting viral proteins. An exemplary virus detection method uses a glycan-based capture matrix (beads) to extract the target virus from the background matrix (e.g., other non-viral biomolecules, contaminants). An aliquot of a sample collected using the sample collection system 1300 may contain other background contaminants and can be coated onto beads carrying capture probes. One or more of glycans, heparin, or carbohydrates can be used as capture materials or probes bound to resin beads or similar types of beads. An optional wash step can be used to remove non-target viral contaminants. The concentrated, purified virus is eluted from the beads using an appropriate solvent and placed in a sealed heating chamber containing organic acids, including formic acid or acetic acid, and heated to 120°C for approximately 10 minutes to degrade protein toxins into specific peptide fragments. This high-temperature, acidic protein digestion protocol cleaves proteins at aspartic acid residues, creating highly reproducible peptide patterns. The above-described capture and digestion processes can be performed using antibodies and enzymes, respectively. This exemplary sample processing, when used with a MALDI-TOFMS, resulted in a sensitivity of over 100 ng / mL for ricin biotoxin (S / N ratio of approximately 50:1) in clean buffer. A 3:1 S / N (signal-to-noise ratio) resulted in a limit of detection (LOD) of less than 10 ng / mL. For a 1 μL sample used in a MALDI-TOFMS analysis system, an LOD of approximately 10 ng / mL corresponds to a total mass on the probe of approximately 10 pg (10-12 g), which corresponds to approximately 20,000 virus particles. An exemplary microfluidic sample processing system for performing the methods disclosed above can be configured to analyze samples collected from the air or other sources, such as nasal swabs. The glycan-based capture column and other microfluidic components are reusable. Large fluid reservoirs containing buffers, weak acids, and alcohols can be used to provide sufficient capacity for measuring hundreds of samples in a single channel of the system. Multiple systems can be run in parallel to process multiple samples simultaneously.The system is cost-effective as it does not require fragile and expensive biomolecular reagents.
[0044] Hot acid digestion reproducibly cleaves proteins at aspartic acid residues, creating known peptide sequences with known masses. These peptide mass distributions are characteristic of precursor proteins. Therefore, digestion provides excellent specificity when the protein of interest is significantly different from background material. Furthermore, the peptide mass distribution is directly determined by the genome, taking post-translational modifications into account. As soon as a new virus is isolated, its sequence is rapidly determined. The RNA sequence of the SARS-CoV-2 virus can be used to accurately predict protein sequences using modern bioinformatics tools (ExPASy Bioinformatics Portal). These proteins can be "digested" in silico (computer-assisted) using bioinformatics tools to generate theoretical peptide maps. Therefore, peptides resulting from SARS-CoV-2 digestion can be predicted and compared with experimental data to generate a specific MALDI-TOF MS signature for the organism. Reports suggest that the major proteins of SARS-CoV are characterized by a nucleocapsid protein of approximately 46 kDa and a spike protein of 139 kDa. Other proteins present in reasonable amounts are the E, M, and N proteins.
[0045] The specificity of target virus detection requires some degree of background removal, especially when the background contains other proteins. In the presence of large amounts of exogenous proteins, peptide maps may be dominated by non-target peptides. As mentioned previously, the use of affinity capture probes for viral toxins based on glycan-modified agarose beads allows for easy cleanup of toxins, even when background proteins and other biomolecules are present in excess. When analyzing breath for viral targets such as SARS-CoV-2, other human proteins in the breath may interfere with the specificity of detection. To ensure the highest specificity, affinity-based sample cleanup is necessary. Viral detection may require bead materials that offer more selective affinity than the glycan-modified beads mentioned above. For example, dextran-based adsorbents may be used to purify viruses, including coronaviruses, but the affinity of these resins for target viruses may not be satisfactory. Alternatively, carbohydrates may be used to purify viruses and proteins, including target viruses such as SARS-CoV and SARS-CoV-2. Additionally, heparin and heparan sulfate may be used as binders to bind to resin beads. Heparin covalently bound to Sepharose beads (GE Healthcare Life Sciences, Heparin Sepharose 6 Fast Flow Affinity Resin, product number 17099801) can be used in place of glycan capture beads. This resin may enable a bead-based affinity capture system for collecting viral particles from exhaled breath. In an exemplary diagnostic system, a breath sample can be drawn through a capture bed in the sample collection system 1300 to collect particles from the patient's breath. The resin beads (bed) can be washed to remove background material. The viral particles adsorbed to the beads can then be eluted using one or more of a high-concentration acid solution, such as about 12.5% acetic acid, about 5% TFA, about 5% formic acid, or about 10% HCl, and sent to a high-temperature acid digestion chamber to generate characteristic peptides.Peptide samples may be mixed with a MALDI matrix and deposited as a substrate suitable for MALDI TOFMS analysis. The sample may be deposited on a suitable substrate or disc pre-coated with a MALDI matrix.
[0046] The report said that analysis of nasal and throat swabs from influenza and COVID-19 patients revealed that approximately 10 3 From 10 10 Less is known about the number of virus particles exhaled by patients. Other reports suggest that influenza patients produce 10 virus particles per breath for about 30 minutes. 4 If the SARS-CoV-2 output is similar to that of influenza, then the number of particles exhaled may be greater than 10 3 From 10 4 Particle output and particle collection efficiency of greater than 99.9% should be sufficient to identify target virus particles in exhaled breath using the exemplary techniques and systems disclosed herein. Detection times using the exemplary systems and methods will be approximately 10 to 20 minutes, including the steps of sample extraction (breathing), sample collection, sample processing (digestion), and analysis using MALDI TOF-MS. This detection time is significantly faster than existing detection systems.
[0047] Exemplary sample processing components can include a thermal acid digestion module or cartridge for autonomously extracting sample from the packed bed column 1301, performing sample cleanup, performing thermal acid digestion, and providing the sample ready for plating onto a MALDI-TOFS sample substrate or disk. The cartridge can be designed to be reusable by adding the ability to flush the cartridge between uses.
[0048] In the exemplary system and method described herein, the length (L) of the packed bed column within the sample capture element 1301 is approximately 3 mm. The nominal inner diameter of the tubing is approximately 7 mm (D). An exemplary packed bed containing approximately 25 mg of C18 resin beads having a nominal particle size (Dp) of approximately 12 μm to 20 μm results in an L / Dp ratio of approximately 150 to 250, with a D / Dp ratio of approximately 350 to approximately 580. These column parameters have been found to prevent undesirable local flow distribution within the bed and ensure that substantially all resin beads are exposed to the aerosol flow through the bed.
[0049] The disclosed exemplary systems and methods can be used to establish a baseline profile of proteins, metabolites, and lipids in exhaled breath, which can then be used to differentiate breath from patients with various diseases, providing a powerful diagnostic tool for disease detection based on the analysis of non-volatile aerosols in exhaled breath.
[0050] The disclosed exemplary systems and methods may also be used for the detection, monitoring, and treatment of diseases other than respiratory and infectious diseases. Chen et al. (2019) described a top-down proteomic strategy for comprehensive identification of cleaved proteins without the use of chemical derivatization, enzymatic manipulation, immunoprecipitation, or other enrichment techniques. Over 1,000 cleaved proteoforms were identified. Tsai et al. (2022) described mass spectrometry-based diagnostic detection of novel coronavirus disease (COVID-19) as a useful alternative to classical PCR-based diagnostics. They used nanoscale liquid chromatography-tandem MS to identify endogenous peptides found in nasal swab saline transport media and identified endogenous peptides and endogenous protease cleavage sites. They reported that SARS-CoV-2 viral peptides were not readily detected and are highly unlikely to contribute to the accuracy of MALDI-based SARS-CoV-2 diagnostics. Lipton et al. (2018) evaluated the association of specific collagen fragments measured in the serum of two independent metastatic breast cancer cohorts and reported that collagen fragments quantified in pretreatment serum were associated with shorter progression-free time and overall survival in two independent cohorts receiving systemic therapy. Ahmed et al. (2005) measured glycated, oxidized, and nitrated protein adducts released by cellular proteolysis using LC-MS / MS to quantitate increased protein damage and the flux of proteolytic products in blood and urine samples from patients with type 1 diabetes. Parchi et al. (1998) examined genomic DNA isolated from frozen tissues of the cerebral cortex, basal ganglia, and cerebellum of patients using SDS-Page electrophoresis and MALDI TOFMS and found that distinct patterns of cleaved prion protein fragments correlated with distinct phenotypes of P102L Gerstmann-Sträussler-Scheinker disease.
[0051] An exemplary method 400 (FIG. 4) for predicting respiratory tract infections (RTIs) and other patient diseases in intubated patients is disclosed. Clinical trial baseline data for diagnosing the presence or absence of RTIs can be obtained by culturing one or more of a sputum sample, an endotracheal tube sample (ET), or a bronchoalveolar lavage fluid (BAL) for each patient in a group of patients, with or without RTIs. In step 401, truncated proteoforms can be identified in the mass spectrum of exhaled aerosol. As described above, the exhaled aerosol can be selectively captured and extracted into one or more liquid collection samples using a packed-bed column. The one or more liquid samples can be analyzed using mass spectrometry to obtain raw mass spectra. In step 402, statistically significant classes of truncated proteoform signatures of respiratory infections can be identified using mass spectral feature selection, including one or more of significance analysis of microarrays (SAM) 403 or t-tests 404. The p-values can be adjusted using the Benjamini-Hochberg method in step 405. In step 406, a multiple logistic regression method may be used to analyze the statistically significant truncated proteoform classes and clinical parameters including age, sex, race, ethnicity, primary diagnosis, medication, sample collection time, microbial identity, white blood cell count, temperature, fraction of inspired oxygen (FiO2) content, lung radiograph, or one or more of the truncated proteoforms within the class to narrow down the statistically significant subset of truncated proteoform classes identified in step 402.
[0052] The presence of an RTI can be predicted using at least one of calculating a composite score representing a statistically significant subset of truncated proteoforms in step 407 and calculating an area under the curve (AUC) of a receiver operating characteristic curve (ROC) representing a statistically significant subset of truncated proteoforms in the sample in step 410.
[0053] Predicting the presence of RTI by calculating a composite score representing a statistically significant subset of truncation proteoforms can include using a reference data sample containing a statistically significant subset of truncation proteoforms to determine a reference threshold mass spectral intensity value (cutoff value in step 409) for each truncation proteoform as a value equal to the normalized mass spectral intensity value (log 10) associated with the intersection of the specificity curve and the sensitivity curve in the ROC for each proteoform (see FIG. 3D). Then, if the measured intensity value of the truncation proteoform is equal to or greater than the reference threshold intensity value, the truncation proteoform is assigned an index score of 1, and if the measured intensity value of the proteoform is less than the reference threshold intensity value, an index score of 0 is assigned. In step 407, the index scores assigned for each statistically significant truncation proteoform in the subset are summed (added) to calculate a composite score representing a statistically significant subset of truncation proteoforms for each sample collected. A cutoff classifier value can be determined that represents the minimum number of statistically significant truncation proteoforms in the subset required to predict the presence of RTI. If the composite score is equal to or greater than the cutoff classifier value, the presence of RTI is predicted. The cutoff classifier value can be determined by generating a confusion matrix for each classifier value, including n, (n-1), (n-2), ..., 0, where n is the total number of statistically significant proteoforms in the subset, using each proteoform's index score (0 or 1) as a predictor and the baseline data as an actual index of RTI (0 or 1). RTI prediction accuracy can be calculated using the confusion matrix for each classifier value, defined as the ratio of the sum of true-positive and true-negative results (TP + TN) to the total number of fluid samples collected (Table 5). The cutoff classifier value can be determined as the classifier value that includes the number of truncated proteoforms required to obtain an RTI prediction accuracy of at least about 90%.
[0054] Identifying statistically significant classes of truncated proteoforms using a t-test can include applying a two-tailed unpaired t-test to the truncated proteoforms in step 404 and adjusting the p-value using the Benjamini-Hochberg method to a 0.05 false discovery rate (FDR) in step 405. The refinement step can include selecting truncated proteoforms with p-values less than 0.05 obtained from the multiple logistic regression analysis to generate a statistically significant subset of truncated proteoforms.
[0055] The exemplary method 400 may further determine whether the composite score is statistically significant in distinguishing between RTI and non-RTI patients if the p-value of the composite score obtained from a multiple logistic regression analysis of variables including one or more of age, sex, race, ethnicity, primary diagnosis, medication, sample collection time, microbial identity, white blood cell count, temperature, fraction of inspired oxygen (FiO2) content, lung radiograph, individual scores of truncated proteoforms within the subset, or the composite score is less than 0.001.
[0056] The presence of an RTI can also be predicted in step 410 (FIG. 5A) by calculating the area under the curve (AUC) of a composite receiver operating characteristic (ROC) curve representing a statistically significant subset of the class of truncated proteoforms. A ROC representing all proteoforms in the statistically significant subset is constructed, and the index scores for each proteoform are used as predictors of RTI, and baseline data are used as actual indicators of RTI to calculate the specificity (TN / TN+FP) and sensitivity (TP / TP+FN) values of the ROC. The area under the curve (AUC) can be determined using a ROC representing all proteoforms in the statistically significant subset. An AUC value of at least about 95% or greater can indicate the presence of an RTI. A statistically significant subset of the class of truncated proteoforms may include one or more of CO6A3 (amino acids 2781-2792), CYTA(2-17), DEN2B(628-637), IRAK4(121-130), MMP9(673-691), or PHTF2(271-285).
[0057] A predictive model for RTI developed using exemplary method 400 can be used to diagnose RTI in a patient. An exemplary method for diagnosing respiratory tract infections (RTIs) in intubated patients by capturing cleaved proteoforms in exhaled aerosols can include selectively capturing cleaved proteoforms in exhaled aerosols produced by each patient using a packed-bed column removably connected to the exhalation tube of a ventilator, extracting the cleaved proteoforms into one or more collected liquid samples corresponding to each patient, analyzing the collected samples corresponding to each patient containing the cleaved proteoforms using mass spectrometry to obtain raw mass spectra, calculating a composite score of statistically significant proteoforms in the samples (where the statistically significant proteoforms are provided by reference data as described above), and diagnosing the presence of RTI if the composite score is equal to or greater than a composite score of the reference data ( FIG. 5A ) that predicts RTI with at least greater than 90% accuracy. A composite score for statistically significant proteoforms in a sample can be calculated by determining the normalized mass spectral intensity value (log10) of each statistically significant truncated proteoform, assigning an index score of 1 to the statistically significant truncated proteoform if its normalized intensity value is equal to or greater than its reference threshold intensity value (Figure 3D), and assigning an index score of 0 to the proteoform if its normalized intensity value is less than its reference threshold intensity value, and adding the index scores to calculate a composite score representing a statistically significant subset of truncated proteoforms in the sample.
[0058] The exemplary systems and methods disclosed above may also be used to predict and diagnose other diseases by capturing truncated proteoforms and other biomarkers in exhaled aerosols. [example]
[0059] Example 1. Capture and analysis of exhaled aerosols from patients diagnosed with COVID-19 using an exemplary packed-bed column connected to a ventilator
[0060] The exemplary system 1300 (Figure 1A) was evaluated in a hospital intensive care unit (ICU) dedicated to treating patients diagnosed with COVID-19 disease. The flow rate through a packed-bed column containing approximately 25 mg of C18 beads (nominal diameter 20 μm) within the sample capture element 1301 was set at 500 ml / min. Before installation in the system 1300, the capture element was washed once with 70% acetonitrile and then three times with 0.05% TFA. The capture element was stored at 4°C prior to use to prevent drying of the C18 beads within the packed bed. Exhaled aerosol was then collected from each patient at a flow rate of 500 ml / min for approximately 4 hours. After the collection period, the packed-bed column was removed from the collection system. The column was washed with approximately 200 μL to approximately 400 μL of 70% ACN or 70% IPA. The organic solvent was removed from the packed-bed column by lyophilization overnight. The organic solvent can also be removed by placing the device 1301 on a heating block at approximately 70°C for approximately 30 minutes. The trapped aerosol particles were extracted or separated using approximately 40 μL to 100 μL of 0.05% TFA. The samples were then analyzed using SDS-PAGE electrophoresis and silver staining, MALDI-TOFMS (whole-cell top-down proteomics), and bottom-up proteomics.
[0061] Approximately 5 μl of each collected sample was used for SDS-PAGE electrophoresis, which was performed using the Criterion Tris-HCl gel system (Bio-Rad Laboratories, Hercules, CA). After SDS-PAGE electrophoresis, SDS-PAGE gels were prepared using a silver staining kit (Thermo Fisher Scientific) to visualize protein bands. Bovine serum albumin was used as an internal positive control. Protein bands were observed in all three patient samples. Based on the BSA control sample, the protein content of the three samples was estimated to be at least 100 ng.
[0062] For whole-cell MALDI-TOF MS analysis, 0.2 μL of analyte was mixed with 0.2 μL of α-cyano-4-hydroxycinnamic acid MALDI matrix (CHCA) prepared in 70% ACN. The mixture was deposited onto a MALDI sample cap, and mass spectra were collected using the exemplary MALDI-TOF mass spectrometry system disclosed in co-pending patent application PCT / US20 / 48042, entitled "System and Method for Rapid and Autonomous Detection of Aerosol Particles," which is incorporated herein by reference in its entirety. MALDI-TOF spectra were collected from samples from patients #3 and #4. Mass peaks were observed in both samples. Peak patterns generated from MALDI-TOF MS were examined using pattern recognition algorithms for detection and classification.
[0063] For bottom-up proteomics, 5 μl of each sample was used. Approximately 50 μl of 50 mM ammonium bicarbonate (pH 8.5) was added to each sample. Protein reduction was performed by adding dithiothreitol to a final concentration of 5 mM and incubating at 37°C for 30 minutes. After reduction, protein alkylation was performed, followed by adding iodoacetamide to a final concentration of 15 mM and incubating at room temperature for 1 hour. Trypsin (Thermo Fisher Scientific) was used for overnight protein digestion. After digestion, peptides were cleaned up using C18 pack tips (Glygen, Columbia, MD). Peptide samples were then prepared in 20 μl of 0.1% formic acid for mass spectrometry analysis, including MALDI-TOF mass spectrometry. Samples were processed using an EASY-nLC 1000 system (Thermo Fisher Scientific) connected to an LTQ Quadrupole-Orbitrap mass spectrometer (Thermo Fisher Scientific). For tandem mass spectrometry, peptides were loaded onto an Acclaim PepMap 100 C18 trap column (0.2 mm x 20 mm, Thermo Fisher Scientific) at a flow rate of 5 μl / min and separated on an EASY-Spray HPLC column (75 μm x 150 mm, Thermo Fisher Scientific). The HPLC gradient was run at a flow rate of 300 nl / min for 60 min using a mobile phase of 5% to 55% (75% acetonitrile and 0.1% formic acid). Mass spectrometry data collection was performed in data-dependent acquisition mode. The resolution for precursor scans was set to 30,000, and the scan resolution for product ions was set to 15,000. Fragmentation of product ions was achieved using higher-energy collision-induced dissociation at 30% of the total energy. Bottom-up proteomics raw data files were processed with MaxQuant Andromeda software (maxquant.org) against the "human" and "SARS-COV-2" protein databases (uniprot.org) according to standard recommendations and instructions.The human protein database contained 20,395 proteins reviewed, and the SARS-CoV-2 protein database contained 13 proteins reviewed. Peptide fingerprints generated from liquid chromatography profiles and digested peptides were identified using LC-MS and MALDI-TOF MS for all three patient samples. A total of 222 proteins were identified in all three patient samples. Most proteins were found to be derived from human blood, indicating active interactions between the lungs and blood. Typical lung proteins and SARS-CoV-2 proteins were identified, as shown in Table 1.
[0064] [Table 1]
[0065] Example 2. Prediction of RTIs by collecting exhaled aerosol samples from intubated patients with respiratory infections.
[0066] Forty-seven exhaled aerosol samples (liquid) were collected from 30 intubated patients in the neurological ICU at Johns Hopkins Hospital. Clinical parameters, including age, sex, race, ethnicity, primary diagnosis, medications, time of sample collection, microbial identification, white blood cell count, temperature, fraction of inspired oxygen (FiO2) testing, and lung radiography, were also collected. Positive respiratory infections were identified based on clinical criteria determined by a physician and if tube samples, including sputum, endotracheal tube samples (ET), or bronchoalveolar lavage fluid (BAL), were culture-positive in the Johns Hopkins clinical laboratory. This clinical data represents the baseline data for the analyses described below.
[0067] The Case Study System 1300 was used to collect exhaled aerosols. The sample collection element contained C18 resin beads with a nominal diameter of approximately 12 μm to approximately 20 μm. The resin beads were packed between two porous polymer frit disks. The inner diameter of the sample collection element was approximately 7 mm. The length of the packed-bed column was approximately 3 mm. One collection element was used for each aerosol sample. The column was connected to a T-joint attached to the exhaust tubing of a mechanical ventilator. The packed bed was rinsed with water before installation in the System 1300. The collection column was connected to a CO2 sensor (Gas Sensing Solutions Ltd, UK) and a mini-diaphragm pump (Parker Hannifin Corporation, Cleveland, Ohio). The pump flow rate was set to a maximum of 0.5 L / min. The CO2 sensor was used to record individual exhaled CO2 levels in the exhaust tubing of the ventilator. After sample collection, the column was disinfected (decontaminated). The column was then eluted with approximately 300 μL of 70% isopropyl alcohol (IPA) to extract proteins and peptides. The solvent was then removed by overnight lyophilization. After lyophilization, approximately 20 μL to 50 μL of 0.05% TFA was added to each sample, and LC-MS / MS analysis was performed.
[0068] For LC-MS analysis, approximately 18 μL of each sample was injected onto a microflow C18 column (Acclaim™ PepMap™ 100, 75 μm x 2 μm x 250 mm, Thermo Fisher Scientific), and proteins were separated using a gradient of solvent (80% acetonitrile, 0.1% formic acid) from 5% to 70% over 60 min using an EASY-nLC 1000 system (Thermo Fisher Scientific). Ion fragmentation was performed using collision-induced dissociation (CID, 35% collision energy) on an LTQ Orbitrap mass spectrometer (Thermo Fisher Scientific) at a mass resolution of 60,000. Raw mass spectrometry data files were searched against the Human Swiss-Prot protein database, which contains 20,387 reviewed entries, and truncated proteoforms were identified using MaxQuant software (Max-Planck-Institute of Biochemistry).
[0069] The workflow for identifying statistically significant classes of features (cleavage proteoforms) (between non-RTI and RTI samples) (Figure 4) included mass spectrometry data processing, feature selection and ranking to identify classes of cleavage proteoforms statistically significant for RTI prediction, and multiple logistic regression to predict the subset of classes of cleavage proteoforms statistically significant for RTI prediction. Generally, mass spectrometry features were normalized by total ion chromatography. Data transformation, scaling, and centering were performed using the logarithmic transformation method. Significance Analysis of Microarrays (SAM), an omics ranking algorithm, was used to select the most important features contributing to distinguishing patients with RTI from those without RTI. SAM provides feature rankings based on the statistics and fold change of each feature (Figures 3B and 4). A two-tailed unpaired t-test was also applied to RTI and non-RTI patients, and for all cleavage proteoforms identified in this study, and raw p-values were obtained. Subsequently, adjustment was made using the Benjamini-Hochberg method, applying a false discovery rate (FDR) of 0.05.
[0070] Multiple logistic regression analysis 406 was used to evaluate correlations between variables, including the patient's RTI status and the measured clinical parameters, and the classes of cleavage proteoforms with statistical significance identified in step 402. Receiver operating characteristic curves (ROC) were constructed, and the area under the curve (AUC) was calculated for the subset of statistically significant features (cleavage proteoforms) between the RTI and non-RTI groups after p-value adjustment. As described above, cutoff values for the subset of statistically significant cleavage proteoforms were generated based on the specificity and sensitivity values of the respective ROC curves. (Figure 3D)
[0071] A total of 263 cleaved proteoforms of 80 proteins were identified (Table 2). The identified proteins overlapped well with proteins from human exhaled aerosol and BAL proteomes, including blood proteins, lung structural proteins, and cytokines, including blood hemoglobin subunits, S100-A9, S100-A12, albumin, zinc alpha-2 glycoprotein, zinc finger homeobox protein 4, uteroglobin, alpha-actinin 1, desmoglein 1, filamin A, mucin 5B, mucin 19, interleukin-1 receptor-associated kinase 4, and matrix metalloproteinase 9. The distribution of cleaved proteoforms in each sample indicated a higher number of cleaved proteoforms in samples from intubated patients with RTI. Furthermore, the difference in the number of cleaved proteoforms identified in exhaled aerosol samples from RTI and non-RTI patients was statistically significant (Figure 3A). In the box-and-whisker plot (Figure 3A), the boxes indicate quartiles, and the horizontal line within each box indicates the median for each case. The whiskers associated with each "box" indicate the maximum and minimum values for each range. The average number for each case is indicated by a cross. Approximately 125 cleaved proteoforms were identified in samples collected from intubated patients with RTI. The number of cleaved proteoforms in intubated patients without RTI was approximately 55 (Figure 3A).
[0072] [Table 2]
[0073] SAM analysis and the Benjamini-Hochberg method were used to identify statistically significant classes of truncated proteoforms that contribute to the separation of RTI and non-RTI samples. Both methods provide statistical significance analysis using a false discovery rate (FDR) adjustment (p = 0.05) for feature reduction. SAM analysis provides feature importance rankings based on their separating power between RTI and non-RTI samples. Six truncated proteoforms, CO6A3 (amino acids 2781-2792), MMP9 (673-691), PHTF2 (271-285), IRAK (121-130), CYTA (2-17), and DEN2B (628-637), were found to be statistically significantly different between the two groups (Figure 3B). SAM analysis ranked all six truncated proteoforms and identified proteoform CO6A3 as a significant feature (truncated proteoform) in the list (Figure 3B). The distribution of the six truncated proteoforms is shown in Figure 3C and Table 3. As shown in Figure 3C, the mass spectral ion intensity of each truncated proteoform in samples from RTI patients was significantly higher than that of non-RTI patients, except for the proteoform corresponding to the protein PHTF2. After FDR adjustment at p = 0.05, all six truncated proteoforms showed statistical significance between the two groups (Table 3). These six proteoforms are characteristic of respiratory infections (e.g., pneumonia, empyema) caused by various bacteria and fungi, including Pseudomonas aeruginosa, Klebsiella pneumoniae, Citrobacter koseri, methicillin-resistant (MRSA) Staphylococcus aureus, ESKAPE, Enterococcus faecium, Acinetobacter baumannii, and Enterobacter spp. In Figure 3D, the x-axis intensity values (arbitrary units generated by the mass spectrometer) of each identified proteoform are extracted from the MaxQuant search results. Each sample was first normalized using the total intensity value. Missing values (zero values) were replaced with 1000, which is two orders of magnitude lower than the lowest intensity value observed in the sample. Values were then logarithmically transformed to the base 10. For example, a value of 10000 becomes 5 after data transformation.To evaluate the use of the six cleaved proteoforms in determining RTI, we performed multiple logistic regression using patient clinical parameters (age, sex, white blood cell count, temperature, and inspired oxygen content) and variables including the identified proteoforms. A multiple logistic regression model was constructed using the glm() function in RStudio, including the cleaved proteoforms as predictors.
[0074] [Table 3] [Table 4]
[0075] RStudio is an integrated development environment for the R programming language for statistical computing and graphics. GLM in R supports non-normal distributions, accepts various parameters, and can be implemented in R through the glm() function, allowing users to apply various regression models. Three truncated proteoforms, CO6A3, MMP9, and PHTF2, were identified as a statistically significant subset of the proteoform class. These three proteoforms were significantly correlated with the presence of RTI (Model 1, Table 4). The most significant truncated proteoform was found to be MMP9, with a p-value of 0.006 (Table 4). The "variables" in Table 4 refer to factors included in the multiple logistic regression analysis. The variables in Model 1 included patient clinical parameters and the six truncated proteoforms.
[0076] Next, we investigated the accuracy of predicting RTI using one or more proteoforms within a statistically significant subset of the proteoform class. Baseline data from a clinical trial containing 47 exhaled aerosol samples were used as the actual indicator of RTI. The presence or absence of RTI was predicted for each classifier value, including n, (n-1), (n-2), …, 0, where n is the total number of statistically significant proteoforms in the subset. In this example, n = 3. Next, confusion matrices (Table 5) were generated for each classifier value: 0, 1, 2, and 3. The confusion matrix for n = 2 resulted in an accuracy of 93.6% (TN + TP / 47) and an accuracy of 95.8% (TP / TP + FP). The prediction accuracies were 53.2%, 78.7%, 93.6%, and 70.2% for n = 0, 1, 2, and 3, respectively. The prediction accuracy was 53.2%, 71.4%, 95.8%, and 100% for n = 0, 1, 2, and 3, respectively. Because the prediction accuracy exceeded 90%, the cutoff classifier value was set to n = 2. Next, the composite score in step 407 was calculated using mass analysis of reference samples. First, using reference samples containing each of the statistically significant subsets of truncated proteoforms CO6A3, MMP9, and PHTF2, a reference threshold mass spectral intensity value was determined as a value equal to the normalized mass spectral intensity value (log10) associated with the intersection of the specificity and sensitivity curves in the ROC for each proteoform (Figure 3D). Next, the collected liquid samples were used to determine the measured mass spectral intensity value for each statistically significant truncated proteoform within the subset. For each analyzed liquid sample, if the measured intensity value of a truncated proteoform within the subset was equal to or greater than its reference threshold intensity value, that truncated proteoform was assigned a score of 1. If the measured intensity value of a proteoform in a collected fluid sample was less than the reference threshold intensity value, it was assigned a score of 0. For each fluid sample, the individual scores assigned to each truncated proteoform in the subset were added to calculate a composite score representing a statistically significant subset of truncated proteoforms in the collected fluid sample. In this example, the minimum composite score would be 0 and the maximum would be 3. [Table 5]
[0077] RTI can be predicted by determining whether the composite score calculated as described above is equal to or greater than the cutoff classification value. In this example, a composite score of 3 is a strong indicator of the presence of RTI. Using 47 fluid samples, the probability of RTI prediction based on the composite score was shown to be 18%, 92%, and 100% (Figure 5A). Of the 47 samples, 12 had a score of 0, meaning all 12 samples were non-RTI, resulting in a 0% probability of predicting RTI. Eleven of the 47 samples had a score of 1, meaning two of the 11 samples were RTI-positive, resulting in an 18% probability of predicting RTI. Thirteen of the 47 samples had a score of 2, meaning 12 of the 13 samples were RTI-positive, resulting in a 92% probability of predicting RTI. Eleven of the 47 samples had a score of 3, meaning all 11 samples were RTI-positive, resulting in a 100% probability of predicting RTI.
[0078] In Table 4, Score 1, Score 2, and Score 3 represent the scores calculated from the individual cleavage proteoforms CO6A3, MMP9, and PHTF2, respectively, and are equal to 1 in each case. Table 4 shows that when examined using multiple logistic regression analysis (Model 2), the individual scores were not statistically significant in distinguishing between RTI and non-RTI patients. However, the composite score was found to be statistically significant with a p-value of less than 0.001. In step 410, we also examined the ability of the three statistically significant proteoforms, CO6A3, MMP9, and PHTF2, to distinguish between RTI and non-RTI patients using AUC (area under the receiver operating characteristic curve) values. The AUC values (Figure 5B) suggest that the individual cleavage proteoforms may not be useful in distinguishing between RTI and non-RTI patients. The AUC values for the truncated proteoforms of CO6A3, MMP9, and PHTF2 were 88.5%, 79.3%, and 76.5%, respectively, which may not be useful for classifying RTI and non-RTI patients. When a linear regression model was constructed using multiple logistic regression with all three truncated proteoforms, the AUC was found to be 96.9% (Figure 5C). This high AUC value suggests that these three truncated proteoforms, when used together, can be used as a criterion for distinguishing between RTI and non-RTI patients, confirming or complementing predictions made using the combined score calculation. The AUC of a good model was approximately 1, indicating good separation between RTI and non-RTI patients.
[0079] The disclosed exemplary methods and systems can also be used to capture cleaved proteoforms in exhaled breath and ambient air collected using a patient-worn mask in an outpatient setting for active case finding or other diagnostic purposes, as disclosed in commonly owned International Application No. PCT / US22 / 22964, the entire contents of which are incorporated herein by reference.
[0080] While this disclosure has been described in connection with preferred modes of practicing it, those skilled in the art will recognize that many modifications can be made thereto without departing from the spirit of this disclosure. Accordingly, it is not intended that the scope of this disclosure be limited by the foregoing description.
[0081] It is also understood that various modifications can be made without departing from the essence of this disclosure. Such modifications are implicitly included in the description and remain within the scope of this disclosure. It is understood that this disclosure is intended to result in patents covering many aspects of the disclosure, both individually and as a system as a whole, and in method and apparatus modes.
[0082] Moreover, each of the various elements of this disclosure and claims may also be achieved in a variety of ways, and this disclosure should be understood to encompass each such variation, whether it be any device implementation variation, method or process implementation, or simply a variation of any of these elements.
[0083] In particular, it should be understood that each element term may be expressed by equivalent apparatus or method terms, even if only the function or result is the same. Such equivalent, broader, or more general terms should be considered included in the description of each element or operation. Such terms may be substituted where necessary to make clear the implicitly broader scope to which this disclosure is entitled. It should be understood that all operations may be expressed as a means for taking that operation or as an element that causes that operation. Similarly, each physical element disclosed should be understood to encompass a disclosure of the operation that the physical element facilitates.
[0084] Furthermore, for each term used, unless its usage in this application is inconsistent with such interpretation, the common dictionary definition contained, for example, in at least one standard technical dictionary recognized by artisans and the most recent edition of Random House Webster's Unabridged Dictionary, should be understood to be incorporated herein for each term and all definitions, alternative terms, and synonyms.
[0085] Additionally, the use of the transitional phrase "comprising" is used to maintain the "open-ended" claims herein in accordance with conventional claim interpretation. Thus, unless the context requires otherwise, "comprising" is intended to mean the inclusion of a recited element or step or group of elements or steps, but not the exclusion of other elements or steps or group of elements or steps. Such terms should be interpreted in the broadest manner to afford applicant the broadest scope legally permissible.
[0086] [References] 1. N. Ahmed, R. Babaei-Jadidi, SK Howell, PJ Beisswenger & PJ Thornalley, "Degradation products of proteins damaged by glycation, oxidation and nitration in clinical type 1 diabetes," Diabetologia 48, 1590-1603 (2005). 2. Dapeng Chen, Lucia Geis-Asteggiante, Fabio P. Gomes, Suzanne Ostrand-Rosenberg, and Catherine Fenselau, "Top-Down Proteomic Characterization of Truncated Proteoforms," J. Proteome Res. 2019, 18, 11, 4013-4019. 3. Hunt, J., "Exhaled breath condensate: An evolving tool for noninvasive evaluation of lung disease," J. Allergy Clin. Immunol. 2002; 110:28-34. 4. Allan Lipton, Kim Leitzel, Suhail M. Ali, Hyma V. Polimera, Vinod Nagabhairu, Eric Marks, Angelique E. Richardson, Laura Krecko, Ayesha Ali, Wolfgang Koestler, Francisco J. Esteva, Diana J. Leeming, Morten A. Karsdal, Nicholas Willumsen, "High turnover of extracellular matrix reflected by specific protein fragments measured in serum is associated with poor outcomes in two metastatic breast cancer cohorts," Intl. J. Cancer, 2018, 43 (11), 3027-3034. 5. J. Brennan McNeil, Ciara M. Shaver, V. Eric Kerchberger, Derek W. Russell, Brandon S. Grove, Melissa A. Warren, Nancy E. Wickersham, Lorraine B. Ware, W. Hayes McDonald, and Julie A. Bastarache, "Novel Method for Noninvasive Sampling of the Distal Airspace in Acute Respiratory Distress Syndrome," American J. Respiratory and Critical Care Medicine 197(8), April 15, 2018. 6. Piero Parchi, Shu G. Chen, Paul Brown, Wenquan Zou, Sabina Capellari, Herbert Budka, Johannes Hainfellner, Patricio F. Reyes, Gregory T. Golden, Jean J. Hauw, D. Carleton Gajdusek, and Pierluigi Gambetti, "Different patterns of truncated prion protein fragments correlate with distinct phenotypes in P102L Gerstmann-Straussler-Scheinker disease," Neuroscience, 95 (14), 8322-8327 (1998). 7. Benjamin Patterson, Carl Morrow, Vinayak Singh, Atica Moosa, Melitta Gqada, Jeremy Woodward, Valerie Mizrahi, Wayne Bryden, Charles Call, Shwetak Patel, Digby Warner, Robin Wood, "Detection of Mycobacterium tuberculosis bacilli in bio-aerosols from untreated TB patients," Gates Open Research 2018, 1:11. 8. Joerg Reifart, Christoph Liebetrau, Christian Troidl, Katharina Madlener and Andreas Rolf, "Noninvasive sampling of the distal airspace via HME-flter fuid is not useful to detect SARS-CoV-2 in intubated patients," Crit. Care (2021) 25:126. 9. Helen Tsai, Brett S. Phinney, Gabriela Grigorean, Michelle R. Salemi, Hooman H. Rashidi, John Pepper, and Nam K. Tran, "Identification of Endogenous Peptides in Nasal Swab Transport Media used in MALDI-TOF-MS Based COVID-19 Screening," ACS Omega 2022, 7, 20, 17462-17471.
Claims
1. 1. A method for predicting respiratory tract infections (RTI) in an intubated patient breathing with ventilator assistance, comprising: For each patient in a group participating in a clinical study, including patients with and without the RTI, diagnosing the presence or absence of the RTI by culturing one or more of a sputum sample, an endotracheal tube sample (ET), or a bronchoalveolar lavage fluid (BAL) to obtain baseline data; selectively capturing truncated proteoforms in exhaled aerosols produced by each patient using a packed bed column removably connected to the exhalation tube of the ventilator; extracting the truncated proteoforms from the packed bed column and providing them in one or more collected fluid samples corresponding to each patient; analyzing said one or more collected liquid samples containing truncated proteoforms using mass spectrometry to obtain raw mass spectra; identifying a statistically significant subset of truncated proteoforms characteristic of said RTI; The method is characterized by comprising a step of predicting the presence of the RTI using one or more of: calculating a composite score representing the statistically significant subset of the truncated proteoforms; or calculating the area under the curve (AUC) of a receiver operating characteristic curve (ROC) representing the statistically significant subset.
2. The step of identifying a statistically significant subset of the truncated proteoforms comprises: referencing the baseline data and using mass spectral feature selection methods including one or more of SAM (Significance Analysis of Microarrays) ranking or t-tests to identify statistically significant classes of truncated proteoforms characteristic of RTIs in the mass spectra; Age, sex, race, ethnicity, primary diagnosis, medication, sample collection time, microbiological identification information, white blood cell count, body temperature, fraction of inspired oxygen (FiO 2 and narrowing down a statistically significant subset of said classes of truncated proteoforms using multiple logistic regression analysis of variables including one or more of: 1) pulmonary sarcoma; 2) pulmonary sarcoma; 3) pulmonary sarcoma; 4) pulmonary sarcoma; and 5) pulmonary sarcoma;
3. 3. The method of claim 2, wherein identifying the class of statistically significant truncated proteoforms using a t-test comprises applying a two-tailed unpaired t-test to the truncated proteoforms and adjusting the p-value using the Benjamini-Hochberg method to apply a false discovery rate (FDR) of 0.
05.
4. 3. The method of claim 2, wherein the narrowing down step comprises selecting truncated proteoforms with a p-value obtained from a multiple logistic regression analysis of less than 0.05 to generate a statistically significant subset of the truncated proteoforms.
5. predicting the presence of the RTI by calculating a composite score representing the statistically significant subset of the truncated proteoforms, using a reference data sample comprising a statistically significant subset of truncated proteoforms, determining a reference threshold mass spectral intensity value for each truncated proteoform in the subset as a value equal to the normalized mass spectral intensity value (loglO) associated with the intersection of the specificity and sensitivity curves in the ROC for each proteoform; assigning an index score of 1 to a truncated proteoform in the subset if the measured mass spectral intensity value (loglO) of the truncated proteoform is greater than or equal to its reference threshold intensity value, and assigning an index score of 0 to a proteoform if the measured mass spectral intensity value of the proteoform is less than its reference threshold intensity value; determining a cutoff classifier value representing the minimum number of statistically significant truncated proteoforms in said subset to predict the presence of an RTI; adding the index scores assigned to each statistically significant truncated proteoform in the subset to calculate a composite score representing a statistically significant subset of the truncated proteoforms for each sample collected; and predicting the presence of an RTI if the composite score is greater than or equal to a cutoff classifier value.
6. determining said cutoff classifier value generating a confusion matrix for each classifier value, including n, (n-1), (n-2), ..., 0, where n is the total number of statistically significant proteoforms in the subset, using each proteoform's index score (0 or 1) as a predictive index and the baseline data as an actual index of RTI (0 or 1); Calculating an RTI prediction accuracy for each classifier value, defined as the ratio of the sum of true positive and true negative results to the total number of fluid samples collected, using the confusion matrix; and determining a cutoff classifier value as a classifier value that includes the number of truncated proteoforms required to obtain an RTI prediction accuracy of at least about 90%.
7. Age, sex, race, ethnicity, primary diagnosis, medication, time of sample collection, microbiological identification, white blood cell count, body temperature, fraction of inspired oxygen (FiO 2 6. The method of claim 5, further comprising determining whether the composite score is statistically significant in distinguishing between RTI and non-RTI patients when the p-value of the composite score obtained from a multiple logistic regression analysis of variables including one or more of: RTI content, lung radiographs, individual scores of truncated proteoforms within the subset, or the composite score is less than 0.
001.
8. further comprising predicting the presence of an RTI by calculating the area under the curve (AUC) of a receiver operating characteristic curve (ROC) representing all proteoforms within a statistically significant subset of truncated proteoforms; This step is constructing a ROC representing all proteoforms in the statistically significant subset, and calculating specificity and sensitivity values for the ROC using the index score for each proteoform as a predictor of RTI and the baseline data as an actual indicator of RTI; Determining the area under the curve (AUC) using the ROC that represents all proteoforms in the statistically significant subset; and predicting the presence of an RTI if the AUC value is greater than at least about 95%.
9. 1. A method for diagnosing respiratory tract infections (RTI) in intubated patients by capturing truncated proteoforms in exhaled aerosols, comprising: selectively capturing truncated proteoforms in exhaled aerosols produced by each patient using a packed bed column removably connected to the exhalation tube of the ventilator; extracting said truncated proteoforms into one or more collected fluid samples corresponding to each patient; analyzing the collected samples corresponding to each patient containing truncated proteoforms using mass spectrometry to obtain raw mass spectra; calculating a composite score of statistically significant proteoforms in said sample, said statistically significant proteoforms being provided by said reference data of claim 5; diagnosing the presence of an RTI if said composite score is greater than or equal to a composite score in said reference data that predicts an RTI with at least greater than or equal to 90% accuracy.
10. Calculating a composite score for statistically significant proteoforms in the sample comprises: determining the normalized mass spectral intensity value (loglO) of each statistically significant truncated proteoform; assigning an index score of 1 to a statistically significant truncated proteoform if the normalized intensity value of the truncated proteoform is equal to or greater than its reference threshold intensity value, and assigning an index score of 0 to a proteoform if the normalized intensity value of the proteoform is less than its reference threshold intensity value; and adding the index scores to calculate a composite score representing a statistically significant subset of truncated proteoforms in the sample.
11. 2. The method of claim 1, wherein the packed bed column comprises one or more of resin beads having C18 functional groups on their surfaces, cellulose beads having sulfate ester functional groups on their surfaces, or a mixture thereof.
12. 10. The method of claim 1, wherein the resin beads and the cellulose beads have a nominal diameter of at least about 20 μm.
13. 10. The method of claim 1, wherein the resin beads and the cellulose beads have a nominal diameter of about 40 μm to about 150 μm.
14. 2. The method of claim 1, wherein extracting the cleaved proteoforms comprises flushing the packed bed column with one or more solvents and collecting the solvent containing the cleaved proteoforms from the packed bed.
15. 15. The method of claim 14, wherein the one or more solvents include one or more of acetonitrile, methanol, trifluoroacetic acid (TFA), or isopropanol (IPA), with the balance being water.
16. 15. The method of claim 14, wherein the one or more solvents comprise about 50% to about 70% by volume of acetonitrile in water, about 50% to about 70% by volume of isopropanol in water, or about 0.05% by volume of TFA in water.
17. 2. The method of claim 1, wherein the statistically significant subset of the class of truncated proteoforms comprises one or more of CO6A3 (amino acids 2781-2792), CYTA (2-17), DEN2B (628-637), IRAK4 (121-130), MMP9 (673-691), or PHTF2 (271-285).
18. In an exhaled breath collection system that captures cleaved proteoforms in exhaled breath aerosols for the diagnosis and treatment of disease, one or more sample capture elements comprising a packed bed column for selectively capturing aerosolized truncated proteoforms in patient-produced breath; A breath collection system comprising: a subsystem configured to be fluidly and electrically coupled to the sample capture element using a quick connect / disconnect coupling, the subsystem including one or more pumps for drawing the exhaled aerosol into the sample capture element, a power source, or a controller for controlling the operation of the sample capture element.
19. 20. The breath collection system of claim 18, wherein the one or more sample capture elements are removably connected to an expiratory tube of a ventilator used to assist breathing in an intubated patient.
20. 20. The breath collection system of claim 18, wherein the controller is configured to detect mechanical and electrical contact between the sample capture element and the subsystem and alert a user via one or more of a graphical user interface and an audible alarm located on the subsystem.
21. The subsystem includes a CO 2 detector disposed between the sample capture element and the pump. 2 20. The breath collection system of claim 18, further comprising one or more of a sensor or a particle counter.
22. 20. The breath collection system of claim 18, wherein the subsystem further comprises a trap disposed between the one or more sample capture elements and the pump and configured to capture exhaled breath condensate (EBC) containing one or more of water vapor, volatile organic components, or non-volatile organic components that passes through the packed bed.
23. The packed bed column may be made of resin, cellulose, silica, agarose, or hydrated Fe 3 O 4 20. The breath collection system of claim 18, comprising solid particles comprising one or more of nanoparticles.
24. 20. The breath collection system of claim 18, wherein the packed bed column comprises one or more of resin beads having C18 functional groups on their surfaces, cellulose beads having sulfate ester functional groups on their surfaces, or a mixture thereof.
25. 20. The breath collection system of claim 18, wherein the resin beads and the cellulose beads have a nominal diameter of at least about 20 μm.
26. 19. The breath collection system of claim 18, wherein the resin beads and the cellulose beads have a nominal diameter of about 40 μm to about 150 μm.
27. 20. The breath collection system of claim 18, wherein the resin beads are packed between two porous polymer frit disks.
28. 19. The breath collection system of claim 18, wherein the nominal flow rate drawn through the bed using the pump is between about 200 ml / min and about 3 L / min.
29. A system for diagnosing and treating diseases by capturing cleaved proteoforms in exhaled breath, The breath collection system of claim 18; a sample extraction system that extracts the captured truncated proteoforms characteristic of the disease from the packed bed column into one or more liquid samples; and an analytical device for analyzing the truncated proteoforms in the one or more liquid samples.
30. 30. The system of claim 29, wherein the extraction system includes means for flushing the packed bed column with at least one solvent and collecting the solvent containing the cleaved proteoforms from the packed bed.
31. 30. The system of claim 29, wherein the analytical device comprises one or more of a PCR, an ELISA, an rt-PCR, a mass spectrometer (MS), a MALDI-MS, an ESI-MS, a MALDI-TOF MS, or an LC-MS / MS.
32. 1. A method for predicting disease by capturing truncated proteoforms in exhaled aerosols, comprising: diagnosing the presence or absence of a disease for each patient in a group of patients with or without the disease participating in a clinical study by culturing one or more of a sputum sample, an endotracheal tube sample (ET), or a bronchoalveolar lavage fluid (BAL) to obtain baseline data; selectively capturing truncated proteoforms in the exhaled breath aerosols produced by each patient using a packed bed column; extracting the truncated proteoforms from the packed bed column to generate one or more collected fluid samples corresponding to each patient; analyzing one or more collected liquid samples containing the truncated proteoforms using mass spectrometry to obtain raw mass spectra; identifying a statistically significant subset of truncated proteoforms that are characteristic of disease; and predicting the presence of disease using one or more of: calculating a composite score representative of said statistically significant subset of said truncated proteoforms; or calculating an area under the curve (AUC) of a receiver operating characteristic curve (ROC) representative of said statistically significant subset.
33. The step of identifying a statistically significant subset of said truncated proteoforms comprises: Referring to the baseline data, using mass spectral feature selection methods including one or more of SAM (Significance Analysis of Microarrays) ranking or t-tests to identify statistically significant classes of truncated proteoforms in the mass spectra that are characteristic of disease; Age, sex, race, ethnicity, primary diagnosis, medication, sample collection time, microbiological identification information, white blood cell count, body temperature, fraction of inspired oxygen (FiO 2 and narrowing down a statistically significant subset of classes of truncated proteoforms using multiple logistic regression analysis of variables including one or more of: 1) pulmonary sarcoma; 2) pulmonary sarcoma; 3) pulmonary sarcoma; 4) pulmonary sarcoma; and 5) pulmonary sarcoma;