Methods and systems for staging monitoring and treating patients with COPD
A machine learning-based system using a random forest classifier model with biomarkers and clinical data predicts COPD composite endpoints, enhancing personalized treatment recommendations and predictive accuracy beyond traditional methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MT SINAI SCHOOL OF MEDICINE
- Filing Date
- 2025-11-04
- Publication Date
- 2026-05-15
AI Technical Summary
Existing prognostic scores for chronic obstructive pulmonary disease (COPD) are not personalized and fail to capture the spectrum of morbidity and cost, lacking predictive power due to inconsistent performance of individual biomarkers and not accounting for patient heterogeneity.
A system and method utilizing machine learning, specifically a random forest classifier model, processes patient data including demographics, clinical features, and biomarkers like IL-8, CC-16, SP-D, RAGE, CRP, and eotaxin to predict composite endpoints such as long-term oxygen utilization, NIPPV, tracheostomy, or nursing home placement, providing personalized treatment recommendations.
The system accurately predicts clinically meaningful patient-centered outcomes, improving predictive power and personalization in COPD management, outperforming traditional methods like GOLD grading and FEV1 alone.
Smart Images

Figure US2025053902_15052026_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR STAGING MONITORING AND TREATING PATIENTS WITH COPDRELATED APPLICATION
[0001] This application claims priority to U. S. Provisional Patent Application No.63 / 717.166, filed November 6, 2024, titled “Methods for Diagnosing and Treating COPD using Machine Learning Models that track specific diagnostic and disease markers,’7the content of which is incorporated herein by reference in its entirety.FIELD
[0002] The present disclosure relates generally to systems, methods, and non-transitory media for use in diagnosing and treating patients with chronic obstructive pulmonary disease (COPD).BACKGROUND
[0003] Chronic Obstructive Pulmonary Disease (COPD) is a leading cause of morbidity and mortality worldwide. It affects 10% of adults older than 40, is a top-ranked cause of death worldwide, and results in more than $30 billion in healthcare costs annually in the United States alone. Disease progression in COPD is highly variable, with some patients experiencing rapid decline while others maintain stability for years. This heterogeny extends to the molecular level, where distinct imaging, inflammatory, and proteomic profiles have been identified, reflecting multiple biological pathways underlying disease progression.
[0004] Despite this. Global Initiative for Chronic Obstructive Lung Disease (GOLD) guideline driven care is focused on personalizing therapy on only two domains - exacerbation rate and symptom burden, and more recently serum eosinophils have been used as an indication for biologic therapy in exacerbation prone COPD. Multiple protein biomarkers have been investigated for prognostic value in COPD, focusing on outcomes such as diagnosis of COPD, decline in forced expiratory volume in one second (FEV1), exacerbation risk, and mortality. The most studied include systemic inflammatory markers (such as C-reactive protein (CRP) and soluble receptor for advanced glycation end products (s-RAGE)), inflammatory cytokines (such as interleukin-8 (IL-8) and interleukin-6 (IL-6)), proteases / antiproteases (such as Alpha-1 antitrypsin), markers related to lung epithelial inflammation (such as pulmonary surfactant protein D (SP-D) and club cell secretory protein (CC-16)), and markers of eosinophilic inflammation (such as sputum eosinophils and eotaxin).
[0005] While the predictive power of any one biomarker has been inconsistent across cohorts, combinations of multiple biomarkers augmented predictive power. Existing prognostic scores are not personalized and do not capture the spectrum of morbidity and cost in COPD.SUMMARY
[0006] The present disclosure addresses these deficiencies by providing systems, methods, and non-transitory media that use machine learning to process COPD patient data including demographics, clinical features, and specific biomarkers to provide accurate, personalized information concerning COPD patient diagnostic and treatment recommendations. The present disclosure enables the prediction of the development of chronic respiratory failure, defined as long term dependence on oxygen, invasive or non-invasive ventilation, or debility with nursing home placement. This composite endpoint represents a clinically meaningful patient-centered outcome reflecting the transition to end-stage lung disease, a major driver of health care utilization and long-term morbidity, and the identification of candidates for early and intensive therapy.
[0007] In some embodiments, a non-transitory processor-readable medium can store code to be executed by a processor of a first compute device. The code can include code to cause the processor to perform or control performance of operations to receive, from a second compute device remote from the first compute device, a trained random classifier model. The code can include code to cause the processor to perform or control performance of operations to receive biomarker data and clinical data of a human subject. The biomarker data can indicate a level of at least one of the following biomarkers: IL-8, CC-16, SP-D, RAGE, CRP, eotaxin, and Alpha- 1 antitrypsin (Al AT), and ratios to one another of any of the preceding. In some embodiments, the model correlates this information with clinical features that can include: FEV1, body mass index (BMI), height, age, and Global Initiative for Chronic Obstructive Lung Disease (GOLD) Stage at enrollment. The code can include code to cause the processor to perform or control performance of operations to execute the trained random classifier model to generate an output indicating a composite COPD endpoint.
[0008] One aspect of the present disclosure provides a method comprising: receiving, for a human patient that has chronic obstructive pulmonary disease (COPD): a respective set of clinical data for said human patient; and a respective set of biomarker data collected from asample from said human patient; and generating, from the respective data sets, one or more human patient features for input to a random forest classifier model; wherein said random forest classifier model has been trained using features generated from individual human subject data collected from a population of human subjects that have COPD; wherein at least one of said random forest classifier model features is generated from FEV1 and wherein said remaining random forest classifier model features are generated from individual human subject data that includes: (a) a respective set of biomarker data and a respective set of clinical data; and (b) a COPD status of said individual human patient; wherein said human patient features match said random forest classifier model features; and generating, by the random forest classifier model, an output for said human patient indicating a COPD status of said patient, wherein the output indicates a likelihood of one composite COPD endpoint out of four possible composite COPD endpoints; and further comprising providing a recommendation for initiating treatment directed to a particular composite COPD endpoint.
[0009] In certain embodiments of said method, said composite COPD endpoint is chosen from a group comprising: (i) long-term oxygen utilization, (ii) long-term Non-Invasive Positive Pressure Ventilation (NIPPV), (iii) tracheostomy, or (iv) long-term nursing home placement. In certain embodiments, said receiving further comprises a respective set of demographic data for said human patient, and wherein said one or more features for input to the random forest classifier model is generated from said demographic data. In certain embodiments, said random forest classifier model determines whether there is a relationship between: (a) a plurality of features derived from at least: the respective set of biomarker data, the respective set of clinical data, and the respective set of demographic data; and (b) the COPD status of said patient. In certain embodiments, said biomarker data indicates a level of at least: receptor for advanced glycation end products (RAGE) and of club cell secretory protein- 16 (CC-16). In certain embodiments, said biomarker data further includes a level of at least one of: interleukin-8 (IL-8), pulmonary surfactant protein-D (SP-D), C-reactive protein (CRP), eotaxin, and Alpha- 1 antitrypsin (Al AT). In certain embodiments, said clinical data is selected from a group comprising: Global Initiative for Chronic Obstructive Lung Disease E (GOLD-E) phenotype over a prior year. Body Mass Index (BMI), age, height, long-term oxygen utilization, longterm Non-Invasive Positive Pressure Ventilation (NIPPV), tracheostomy, and long-term nursing home placement. In certain embodiments, said demographic data is selected from a group comprising: Body Mass Index (BMI), height, and age. In certain embodiments, said random forest classifier model has been trained using hyperparameter optimization using grid-search and 10-fold cross-validation of training data, and wherein hyperparameters include number of trees, tree depth, minimum samples per split and leaf, impurity criterion, and number of features considered. In certain embodiments, said method further comprises sequentially removing clinical variables until loss of cross-validation performance on a training set is observed.
[0010] One aspect of the present disclosure provides a system comprising: one or more processors; and one or more non-transitory computer-readable media, coupled to the one or more processors and storing instructions which, when executed by the one or more processors, cause the one or more processors to perform or control performance of operations comprising: receiving, for a human patient that has chronic obstructive pulmonary disease (COPD): a respective set of clinical data for said human patient; and a respective set of biomarker data collected from a sample from said human patient; and generating, from the respective data sets, one or more human patient features for input to a random forest classifier model; wherein said random forest classifier model has been trained using features generated from individual human subject data collected from a population of human subjects that have COPD; wherein at least one of said random forest classifier model features is generated from FEV1 and wherein said remaining random forest classifier model features are generated from individual human subject data that includes: (a) a respective set of biomarker data and a respective set of clinical data; and (b) a COPD status of said individual human patient; wherein said human patient features match said random forest classifier model features; and generating, by the random forest classifier model, an output for said human patient indicating a COPD status of said patient, wherein the output indicates a likelihood of one composite COPD endpoint out of four possible composite COPD endpoints; and further comprising providing a recommendation for initiating treatment directed to a particular composite COPD endpoint.
[0011] In certain embodiments of said system, said composite COPD endpoint is chosen from a group comprising: (i) long-term oxygen utilization, (ii) long-term Non-Invasive Positive Pressure Ventilation (NIPPV), (iii) tracheostomy, or (iv) long-term nursing home placement. In certain embodiments, said receiving further comprises a respective set of demographic data for said human patient, and wherein said one or more features for input to the random forest classifier model is generated from said demographic data. In certain embodiments, said random forest classifier model determines whether there is a relationship between: (a) a plurality of features derived from at least: the respective set of biomarker data, the respectiveset of clinical data, and the respective set of demographic data; and (b) the COPD status of said patient. In certain embodiments, said biomarker data indicates a level of at least: receptor for advanced glycation end products (RAGE) and of club cell secretory protein- 16 (CC-16). In certain embodiments, said biomarker data further includes a level of at least one of: interleukin-8 (IL-8), pulmonary' surfactant protein-D (SP-D), C-reactive protein (CRP), eotaxin, and Alpha- 1 antitrypsin (Al AT). In certain embodiments, said clinical data is selected from a group comprising: Global Initiative for Chronic Obstructive Lung Disease E (GOLD-E) phenotype over a prior year. Body Mass Index (BMI), age, height, long-term oxygen utilization, longterm Non-Invasive Positive Pressure Ventilation (NIPPV), tracheostomy, and long-term nursing home placement. In certain embodiments, said demographic data is selected from a group comprising: Body Mass Index (BMI), height, and age. In certain embodiments, said random forest classifier model has been trained using hyperparameter optimization using gridsearch and 10-fold cross-validation of training data, and wherein hyperparameters include number of trees, tree depth, minimum samples per split and leaf, impurity criterion, and number of features considered. In certain embodiments, said method further comprises sequentially removing clinical variables until loss of cross-validation performance on a training set is observed.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 provides a diagram of patient diagnosis and treatment system, according to an embodiment of the present disclosure.
[0013] Figure 2 provides a diagram of a computing device, according to an embodiment of the present disclosure.
[0014] Figure 3 provides a flowchart showing an individualized method for diagnosing and characterizing patients with COPD, according to an embodiment of the present disclosure.
[0015] Figure 4 provides a flowchart showing an example of the development and use of the methods and systems of an embodiment.
[0016] Figure 5 provides an example study flow diagram, according to an embodiment of the present disclosure.
[0017] Figure 6 provides example receiver operating characteristic (ROC, panel A) and precision-recall (PR curve, panel B) curves comparing model performance for the final model in the training and testing sets, the GOLD grade, and an FEVi-only model evaluated on the testing set.
[0018] Figure 7 provides an example of the SHapley Additive exPlanations (SHAP) summary plot illustrating each feature’s contribution to the predicted probability of the composite respiratory failure outcome; red points represent higher feature values and blue points lower values, with rightward displacement indicating increased predicted risk (e.g., lower FEVi is associated with higher risk).DETAILED DESCRIPTION
[0019] Non-limiting examples of various embodiments of the present disclosure are provided herein and illustrated in the accompanying figures and drawings.
[0020] Certain markers have been associated with COPD, such as, systemic inflammatory markers (such as CRP and s-RAGE), inflammatory cytokines (such as IL-8 and IL-6), proteases / antiproteases (such as alpha- 1 antitry psin), markers related to lung epithelial inflammation (such as SP-D and CC-16), and markers of eosinophilic inflammation (such as sputum eosinophils and eotaxin). Provided in the present disclosure are systems, methods, and non-transitory media that utilize machine learning to screen a cultivated combination of biomarkers, clinical features, and demographic information to identify possible endpoints for COPD patients.
[0021] An illustration of certain embodiments of the models and sy stems is provided in Figure 1, which provides in an implementation consistent with the disclosure. In the figure, the patient diagnostic and treatment system 100 includes a hardware-based processor 102, a memory 104 configured to store instructions and configured to provide the instructions to the hardware-based processor 102, an input / output device 106, and a machine learning model 110 configured to implement the instructions provided to the hardware-based processor 102.The model 110 includes a random forest classifier model. In an implementation, the instructions are code written in any appropriate programming language.
[0022] An implementation of certain methods consistent with the disclosure as provided in Figure 1 provides: receiving subsequent patient data 124 as features at the input output device106 for input into the random forest classifier (RF) model 110 that has been trained previously on training data 122 from patients diagnosed with COPD; and generating, with said RF model that was previously trained on said patient data 122 and based on said subsequent patient data 124, an output 126 for the diagnosis and / or treatment of COPD in said patient.
[0023] In one implementation, the data source 120 is in proximity to the diagnostic system 100. In another implementation, the data source 120 is remote from the diagnostic system 100. In a further implementation, the training data 122 is obtained from a database of patient information. For example, the training data 122 optionally includes data obtained from the Mount Sinai BioMe BioBank. In one implementation, the training data 122 are formatted as structured extensible markup language (XML) files including both raw data and metadata associated with patient identifiers, time, place, indication, and characteristics such as diagnoses of the patients associated with each of the training data 122.
[0024] The biomarker data of the set of human subjects can indicate a level of at least one of the following biomarkers: interleukin-8 (IL-8), club cell secretory protein-16 (CC-16 or uteroglobin), pulmonary' surfactant protein-D (SP-D or PSP-D), receptor for advanced glycation end products (RAGE), C-reactive protein (CRP), eotaxin, Alpha-1 antitrypsin (Al AT), interleukin 5 (IL-5), interleukin 6 (IL-6), interleukin 13 (IL- 13), and interleukin 33 (IL-33), and ratios to one another of any of the preceding. Clinical information can include, for example, body mass index (BMI), Global Initiative for Chronic Obstructive Lung Disease (GOLD)-E phenoty pe in the last year, smoking status, smoking history measurements of activity, height, and age. Demographic information can include, for example, sex, race, ethnicity, age, BMI, and height.
[0025] In one implementation, the patient data 124 is received in real time. Such real time data 124 is temporarily or permanently stored in the data source 120, and is transmitted, conveyed, or otherwise provided to input / output device 106 of the diagnostic and treatment system 100. In another implementation, the patient data / information 124 data are formatted as structured extensible markup language (XML) files including both raw data and metadata associated with patient identifiers, time, place, indication, and additional patient information received from the data source 120. In further implementations, the patient data 124 are formatted in any appropriate data format. In one implementation, the output 126 is an alert, a notification, or a message output from the input / output device 106 containing a treatment recommendation. For example, the input / output device 106 includes a display or monitorconfigured to visually display the treatment recommendation as an output 126 to a doctor, a technician, or a patient. The output 126 is a text message or an image representing the diagnostic state of the patient, such as the patient corresponding to the patient data 124, as requiring treatment for chronic obstructive pulmonary disease (COPD). The output 126 optionally includes an audio signal.
[0026] Figure 2 provides a diagram of a computing device 200, according to an embodiment of the present disclosure. The computing device 200 can be implemented in / by the system 100 of Figure 1 in some embodiments. In embodiments of the present disclosure, as illustrated in Figure 2, the computing device 200 can include different components. In another alternative implementation, the functions of a given component can instead be carried out by one or more multiple different components. The computing device 200 can be implemented by a virtual computing device, using a cloud computing environment, or by a plurality of any known computing devices. In some implementations, the device 200 includes a processor 202, a memory 204, an input / output (I / O) interface and tools 206. and input / output (I / O) devices 214.
[0027] The processor 202 can be a hardware-based processor implementing a system, a sub-system, or a module. The processor 202 can include one or more general-purpose processors. Alternatively, the processor 202 can include one or more special-purpose processors. The processor 202 can be integrated in whole or in part with the memory 204 and the input / output interface 206. In another alternative implementation, the processor 202 can be implemented by any known hardware-based processing device such as a controller, an integrated circuit, a microchip, a central processing unit (CPU), a microprocessor, a system on a chip (SoC), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In addition, the processor 202 can include a plurality of processing elements configured to perform parallel processing. In a further alternative implementation, the processor 202 can include a plurality of nodes or artificial neurons configured as an artificial neural network. The processor 202 can be configured to implement any suitable machine learning (ML) based devices, any suitable artificial intelligence (Al) based devices, and any suitable artificial neural networks, including a recursive neural network (RNN) or a convolutional neural network (CNN). A processor may include a system with a general-purpose central processing unit (CPU), multiple processing units, dedicated circuitry for achieving functionality, or other systems. Processing need not be limited to a particulargeographic location, or have temporal limitations. For example, a processor may perform its functions in “real-time,” “offline.” in a “batch mode,” etc. Portions of processing may be performed at different times and at different locations, by different (or the same) processing systems. A computer may be any processor in communication with a memory. The processor 202 can be implemented as the processor 102 of Figure 1 in some embodiments.
[0028] The memory 204 can be implemented as a non-transitory computer-readable storage medium such as a hard drive, a solid-state drive, an erasable programmable read-only memory (EPROM), a universal serial bus (USB) storage device, a floppy disk, a compact disc read-only memory (CD-ROM) disk, a digital versatile disc (DVD), cloud-based storage, or any known non-volatile storage suitable for storing instructions for execution by the processor, and located separate from processor 202 and / or integrated therewith. The memory 204 can store software executable on the computing device 200 by the processor 202, including an operating system 208, one or more applications 210 and its related data 212. The memory 204 can be implemented as the memory 104 of Figure 1 in some embodiments.
[0029] The code of the processor 202 can be stored in a memory internal to the processor 202. The code can be instructions implemented in hardware. Alternatively, the code can be instructions implemented in software. The instructions can be machine-language instructions executable by the processor 202 to cause the computing device 200 to perform the functions of the computing device 200 described herein. Alternatively, the instructions can include script instructions executable by a script interpreter configured to cause the processor 202 and computing device 200 to execute the instructions specified in the script instructions. In another alternative implementation, the instructions are executable by the processor 202 to cause the computing device 200 to execute an artificial neural network, the model 110 of Figure 1, etc. The processor 202 can be implemented using hardware or software, such as the code. The processor 202 can implement a system, a sub-system, or a module, as described herein.
[0030] The memory 204 can store data in any known format, such as databases, data structures, data lakes, or network parameters of a neural network. The data can be stored in a table, a flat file, data in a filesystem, a heap file, a B+ tree, a hash table, or a hash bucket. The memory 204 can be implemented by any known memory, including random access memory (RAM), cache memory, register memory, or any other known memory device configured to store instructions or data for rapid access by the processor 202, including storageof instructions during execution. Any of software in the memory 204 can alternatively be stored on any other suitable storage location or computer-readable medium. In addition, memory 204 (and / or other connected storage device(s)) can store instructions and data used in the features described herein. The memory 204 and any other type of storage (magnetic disk, optical disk, magnetic tape, or other tangible media) can be considered "storage" or "storage devices."
[0031] The I / O interface 206 can be any known device configured to perform the communication interface functions of the computing device 200 described herein. The I / O interface 206 can implement wired communication between the computing device 200 and another entity. Alternatively, the I / O interface 206 can implement wireless communication between the computing device 200 and another entity. The I / O interface 206 can be implemented by an Ethernet, Wi-Fi, Bluetooth, or USB interface. The I / O interface 206 can transmit and receive data over a network and to other devices using any known communication link or communication protocol. The I / O interface 206 can be implemented as the I / O device 106 of Figure 1 in some embodiments.
[0032] The I / O interface and tools 206 can provide functions to enable interfacing the computing device 200 with other systems and devices. For example, network communication devices, storage devices, and input / output devices can communicate with the computing device 200 via an I / O interface and tools 206. In some implementations, the I / O interface and tools 206 can connect to interface devices including input devices (keyboard, pointing device, touchscreen, microphone, camera, scanner, etc.) and / or output devices (display device, speaker devices, printer, motor, etc.), which are collectively shown as at least one input / output device 214. The tools 206 may further include GPS tools, sensor tools, accelerometers, etc. that may communicate with other components of the computing device 200 or with components external thereto.
[0033] The I / O interface 206 can be any suitable device configured to perform user input and output functions. The I / O interface 206 can be configured to receive an input from a user. Alternatively, the user interface 206 can be configured to output information to the user. The I / O interface 206 can be a computer monitor, a television, a loudspeaker, a computer speaker, or any other known device operatively connected to the computing device 200 and configured to output information to the user. A user input can be received through the user interface 206 implementing a keyboard, a mouse, or any other known device operativelyconnected to the computing device 200 to input information from the user. Alternatively, the I / O interface 206 can be implemented by any known touchscreen. The computing device 200 can include a server, a personal computer, a laptop, a smartphone, or a tablet.
[0034] For ease of illustration, Figure 2 shows one block for each of processor 202, memory 204. I / O interface and tools 206, the application 210, etc. These blocks may represent one or more processors or processing circuitries, operating systems, memories, I / O interfaces, applications, and / or software modules. In other implementations, the device 200 may not have all of the components shown and / or may have other elements including other types of elements instead of, or in addition to, those shown herein.
[0035] A user device can also implement and / or be used with features described herein. Example user devices can be computer devices including some similar components as the computing device 200 (e.g., processor(s) 202, memory 204, and I / O interface and tools 206). An operating system, software and applications suitable for the client device can be provided in memory and used by the processor. The I / O interface for a client device can be connected to network communication devices, as well as to input and output devices, e.g., a microphone for capturing sound, a camera for capturing images or video, audio speaker devices for outputting sound, a display device for outputting images or video, or other output devices. A display device within the audio / video input / output devices 214, for example, can be connected to (or included in) the device 200 to display images pre- and post-processing as described herein, where such display device can include any suitable display device (e.g., an LCD, LED, or plasma display screen, CRT, television, monitor, touchscreen, 3-D display screen, projector, or other visual display device). Some implementations can provide an audio output device (e.g., voice output or synthesis that speaks text).
[0036] One or more methods described herein can be implemented by computer program instructions or code, which can be executed on a computer. For example, the code can be implemented by one or more digital processors (e.g., microprocessors or other processing circuitry), and can be stored on a computer program product including a non-transitory computer readable medium (e.g., storage medium), for example, a magnetic, optical, electromagnetic, or semiconductor storage medium, including semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), flash memory, a rigid magnetic disk, an optical disk, a solid-state memory drive, etc. The program instructions can also be contained in, and provided as, anelectronic signal, for example in the form of software as a service (SaaS) delivered from a server (e.g., a distributed system and / or a cloud computing system). Alternatively, one or more methods can be implemented in hardware (logic gates, etc.), or in a combination of hardware and software. Example hardware can be programmable processors (e.g., field-programmable gate array (FPGA), complex programmable logic device), general purpose processors, graphics processors, application specific integrated circuits (ASICs), and the like. One or more methods can be performed as part of or component of an application running on the system, or as an application or software running in conjunction with other applications and operating system.
[0037] One or more methods described herein can be run in a standalone program that can be run on any type of computing device, a program run on a web browser, a mobile application C’app”) run on a mobile computing device (e g., cell phone, smart phone, tablet computer, wearable device (wristwatch, armband, jewelry, headwear, goggles, glasses, etc.), laptop computer, etc.). In one example, a client / server architecture can be used, for instance, a mobile computing device (as a client device) sends user input data to a server device and receives from the server the final output data for output (e.g., for display). In another example, all computations can be performed within the mobile app (and / or other apps) on the mobile computing device. In another example, computations can be split between the mobile computing device and one or more server devices.
[0038] Figure 3 provides a flowchart show ing an individualized method 300 for diagnosing and characterizing patients w ith COPD, according to an embodiment of the present disclosure. In an implementation of the present disclosure, the patient diagnosis and treatment system 100 can be trained and tested using the method 300 by: receiving data 301 for each training subject (patient) in a set of patients that have COPD; training 302 a random forest classifier using (a) a set of features for each training subject and (b) identification for each training subject a clinical endpoint or status marker for said subject (where the clinical endpoint or status is chosen from a list of possible endpoints); receiving data 303 for a test subject that has COPD but whose data was not included in the training subject data; and generating 304, using the trained random forest classifier of the model 110, an output that indicates the clinical endpoint or status of said test subject. In some embodiments, the training and testing of the method 300 can use data 301 for each training subject or each test subject 303 where the data has been obtained at different time points.
[0039] The system 100 can be configured to determine, for each respective human training subject in the plurality of diabetic human training subjects, a respective indication of a composite COPD endpoint. The random forest classifier model 110 can be configured to be trained based on: (a) a set of features generated from patient biomarker data, and / or a set of features generated from patient clinical data, and, optionally, a set of features generated from patient demographic data, and (b) a respective indication of the composite COPD endpoint for a given patient. The random forest classifier can therefore be iteratively optimized to generate the respective indication of whether a respective human subject is at a particular composite COPD endpoint based on the set of features generated from the biomarker data, and / or the set of features generated from the respective clinical data, and / or the set of features generated from the respective demographic data.
[0040] In some instances, the random forest classifier model 110 can be configured to be trained based on: (a) a set of features generated from patient biomarker data, and / or a set of features generated from patient clinical data, and. optionally, a set of features generated from patient demographic data, and (b) a respective indication of the composite COPD endpoint for a given patient.
[0041] In some embodiments, random forest classifier is trained against, for each respective training subject in the plurality of training subjects, a set of at least four features (e.g., selected from the types of biomarker data, clinical data, and / or demographic data described herein). In some embodiments, the random forest classifier is trained against, for each respective training subject in the plurality of training subjects, a set of at least 5, 6, 7, 8, 9, 10, 11. 12. 13. 14, 15, or more features.
[0042] In some embodiments, the set of features used to train the machine learning model includes at least an abundance value for receptor for advanced glycation end products (RAGE) and of club cell secretory protein- 16 (CC-16). a comparison of (i) an abundance value for RAGE and (ii) an abundance value for CC-16 (e.g., a ratio of the abundance values), a comparison of (i) an abundance value for RAGE and (ii) an abundance value for CC-16 (e.g., a ratio of the abundance values), an interleukin-8 (IL-8) level, an SP-D level, a C-reactive protein (CRP) level, an eotaxin level, and Alpha-1 antitrypsin (Al AT).EXAMPLE 1
[0043] The following provides an example of using the methods of the present disclosure according to an embodiment.Patient Population
[0044] Patients with COPD were identified from the Mount Sinai BioMe Biobank at the Icahn School of Medicine at Mount Sinai from 2011-2022. The BioMe Biobank is an institutional review board-approved biorepository of plasma and DNA collected from 2007 onward in a diverse community in New York City linked to longitudinal Electronic Health Record (EHR) data. Patients with COPD were electronically phenotyped using ICD codes with a previously validated algorithm. Briefly, a COPD diagnosis was established when a patient had 3 or more COPD related ICD codes (491*, 492*, 496*, J41*, J43*, J44*) in a rolling 730.5 day window with age >=35 years old at first diagnostic code. Patients were included if they had banked plasma no more than one year before the first diagnosis of COPD. In patients with multiple available Biobank specimens, only the first eligible sample was enrolled. Patients were excluded if they met any individual endpoint or the composite endpoint at the beginning of follow-up (Figure 5). This study was approved by the Mount Sinai Institutional Review Board and granted a waiver of informed consent (‘STUDY-21-01629’). Patient data used for the study was anonymized to comply with HIPAA regulations.Biomarker Quantification
[0045] Plasma biomarkers were chosen based on previously validated prognostic predictive value in COPD. A multiplex panel of plasma biomarkers associated with COPD outcomes was tested using the Meso Scale Discover}7(MSD) electrochemiluminescence immunoassay (ECLIA) platform. Prior to assay implementation, analytical validation was performed for each biomarker, including assessments of accuracy, precision, linearity, and recovery. All assays were run under strict quality control protocols. In-house control samples were tested in duplicate on ever}7plate to monitor intra- and inter-assay variability. Inter-assay coefficients of variation (CVs) were calculated across multiple runs, with thresholds of <10% for intra-assay and <15% for inter-assay CVs considered acceptable. Levey-Jennings plots were used to track performance over time. This rigorously validated panel enabled robust, reproducible quantification of candidate COPD biomarkers.Outcome Determination and Assessment of the Clinical Endpoint
[0046] Patients were labeled with a binary composite endpoint of (i) long term oxygen utilization, (ii) long term Non-Invasive Positive Pressure Ventilation (NIPPV), (iii) tracheostomy, or (iv) long term nursing home placement. The beginning of follow-up was defined as the later of BioBank plasma selection date or first diagnosis of COPD. Outcome and clinical data were extracted from the electronic health record (EHR) by three trained abstractors with a structured instrument focusing on clearly defined clinical endpoints (Table 3; Figure 4); cases with uncertainty were adjudicated by group consensus. Patient characteristics and demographics were extracted from BioBank enrollment forms. Missing numeric features were imputed using population medians. Categorical data were normalized using an investigator defined mapping and encoded with one-hot-encoding. Numerical scaling and centering were not applied as the chosen model architectures are insensitive to monotonic scaling transformations. Data curation was done in an automated pipeline to maintain reproducibility.Model Development
[0047] The study cohort was randomly split into training (70%) and testing sets (30%). Patient demographics, smoking history, functional status, history of COPD related hospitalizations and exacerbations, spirometry values, and biomarker values were used for model training. Class imbalance in the training set was handled with class-weighting as opposed to majority class under-sampling because of sample size. A random forest model was trained with hyperparameter optimization using grid-search and 10-fold cross-validation of training data (Figure 4). Hyperparameters included number of trees, tree depth, minimum samples per split and leaf, impurity criterion, and number of features considered at each split, with the best model selected based on mean cross-validated accuracy. Iterative feature elimination was performed by sequentially removing clinical variables until loss of cross-validation performance on the training set was observed.Model Validation and Statistical Analysis
[0048] The final tuned model w as evaluated on the held-out testing set. Performance was assessed using the area under the receiver operating curve (AUC), area under the precisionrecall curve (AUPRC). and sensitivity, specificity, accuracy, fl score, and precision at different thresholds. Two comparator models were trained under identical conditions, one including only the GOLD Grade, and the other containing only FEV1 (percent predicted). Differences in AUCwere compared using the method of DeLong. Analyses were performed in Python (version 3.10.14) using scikit-leam (version 1.4.2) for model development and evaluation.Results: Baseline Characteristics of Study Cohort:
[0049] Baseline characteristics of the cohort (n = 931) are as follows (Figure 5, Table 1): mean age 65 years old, 48% male, 83% of patients reported ever smoking more than 100 cigarettes. Median FEV1 was 67 (IQR 54, 81); by GOLD Grade 17% of patients were GOLD grade 1, 37% GOLD 2, 10% GOLD 3, and 1% GOLD 4. 13% of patients had a GOLD E phenotype (2 moderate COPD exacerbations or 1 requiring hospitalization) in the last year. 212 patients (23%) were label positive for composite endpoint: 161 (17.2%) for long-term oxygen use, 54 (5.8%) long-term NIPPV use, 20 (2%) tracheostomy placement, and 54 (5.8%) for nursing-home placement. Patients meeting the composite endpoint had lower FEV 1 (median 56% vs 70% percent predicted) than those who did not. Patients who met the composite endpoint were more likely to be GOLD Grade 3 (24% vs 6.1%) and to have a GOLD E exacerbation-prone phenotype (22% vs 9.9%) than those who did not. Patients meeting endpoint were also had higher rates of congestive heart failure (33% vs 22%); however, other demographic characteristics were similar between groups. The cohort was partitioned into a Training set (n = 652) for model training and optimization and a held-out Testing set (n = 279) for validation.Results: Prediction of the Respiratory7Failure Composite Endpoint:
[0050] A random forest classifier was trained and optimized on the training set (n = 652) and validated on the held-out testing set (n = 279) where comparison was made to models containing only the GOLD Grade and the FEV1 alone (Figure 6). The final model had an AUC-ROC of 0.70 and an AU-PRC of 0.46 on the testing set. Model AUC was higher than the GOLD Grade (AUC 0.59, AAUC=0.10. p<0.01) and a model comprising of FEV1 alone (AUC 0.63, AAUC=0.073, p<0.05). At Fl optimized threshold, the model had sensitivity of 52% and specificity of 80% for the composite outcome with Fl score of 0.48. PPV was 44% and NPV was 85%. Initial model features are provided in Table 4. Final hyperparameters optimized via grid-search are detailed in Table 5.Results: Feature Elimination and Feature Importance:
[0051] After iterative feature elimination, 12 predictors were retained in the final model (Table 2, Table 5). FEV1 (% Predicted) had the highest relative contribution followed by RAGE, CC-16, IL-8, SP-D, CRP, Eotaxin, age, BMI, GOLD E phenotype in the last year, and patient height. Biomarkers collectively contributed 65% of Gini importance. Shapley Additive exPlanations (SHAP) values (Figure 7) indicated that lower FEV1, a GOLD E phenoty pe, and higher BMI were associated with the composite endpoint. Older age was inversely associated with the composite endpoint in the model. For biomarkers, lower levels of RAGE and Al AT and higher levels of IL-8, Eotaxin, CRP, SP-D, and CC-1 were associated with the composite endpoint.Discussion:
[0052] The present disclosure provides an internally validated a random forest derived risk score combining multiple biomarkers, spirometry, and clinical data to predict a composite outcome of chronic respiratory failure in patients with COPD. The final model demonstrated significantly higher predictive power in a held-out test set than the GOLD Grade. The final model had 18.6% improved AUC over the GOLD grade, at F 1 optimized threshold 57% higher precision than the GOLD grade. While baseline FEV1 was, as expected, the single most informative feature, the biomarkers collectively accounted for most of the model predictive power (Gini 0.65). This highlights the additive predictive value of biomarker profiling beyond standard clinical indices, also reflected in that the performance of the model is significantly better than a similarly trained model using FEV1 alone.
[0053] Model interpretability analysis by feature importance with Gini coefficient and SHAP values underscores the functionality of the model in the present disclosure. FEV1 has been consistently demonstrated to be an important feature in previous work combining clincial and biomarker data. Expected relationships of inflammatory markers, such as CRP, sRAGE, and IL8, and lung proteins, such as SP-D, were also observed in the present cohort. SHAP plots suggest a subset of patients with higher age with lower risk of the composite outcome. Lower levels of CC-16 have been associated with increased disease progression in COPD; however, in this cohort higher CC-16 pushed model predictions toward meeting composite endpoint. Two previous studies have assessed the predictive value of CC-16 in combination with other markers: one study demonstrated an improved predictive power when adding CC16 for a variety of endpoints, though the direction w as not analyzed. In another study CC-16 was not significantly associated with mortality' in a multivariate analysis.
[0054] While previous studies have demonstrated that multiple biomarkers can predict outcomes in COPD, these previous studies have been limited by their primarily investigational nature assessing statistical associations as opposed to evaluating the models as a prognostic risk-prediction index. Furthermore, the current study uses a composite end-point, comprised of long term dependence on oxygen, non-invasive positive pressure ventilation, tracheostomy placement, or nursing home placement. The use of these kinds of patient centered outcomes is increasingly preferred in COPD to capture the heterogenous disease progression and clinical phenotypes not captured by traditional measures such as decline in FEV1 and move care beyond preventing exacerbations to altering disease trajectory'.
[0055] The current disclosure provides a deeply phenotyped biobank cohort linked to longitudinal EHR data, which allows for robust characterization of patient centered endpoints. The biomarker panel was developed on a mature platform for high throughput multiplex biomarker quantification, which enables direct translation into clinical use at scale, as has been done in other domains such as diabetic kidney disease. The input features have strong biological rationale, and are straightforward to collect comprising of spirometry; biomarkers available in a single blood draw, and basic demographics - this will further facilitate translation into the clinical workflow in the future.
[0056] Table 1. Characteristics of Patients in Analytic Cohort
[0058] Table 2. Final Model Input Variables and Relative Variable Importance.
[0059] Table 3. Chart Abstraction Questionnaire
[0060] Table 4. Full Input Feature List Before Feature Reductioni numbercigarettess: Number of current cigarettes smoked daily
[0061] Table 5. Final Random Forest Model Hyperparameters. Optimal model hyperparameters were chosen based on grid search of hyperparameter options using 10-fold cross validation on model training data.1 Hyperparame i Description j criterion og loss Splitting critermax depth 4 i Maximum tree depth s max feat ures sqrt Features considered per split jmin samplcs 2 Minimum sampl i min samplcs 2 i Minimum sampl n estimators 100 Number of trees
Claims
CLAIMSWhat is claimed is:
1. A method comprising:receiving, for a human patient that has chronic obstructive pulmonary disease (COPD):a respective set of clinical data for said human patient; anda respective set of biomarker data collected from a sample from said human patient; andgenerating, from the respective data sets, one or more human patient features for input to a random forest classifier model;wherein said random forest classifier model has been trained using features generated from individual human subject data collected from a population of human subjects that have COPD;wherein at least one of said random forest classifier model features is generated from FEV1 and wherein said remaining random forest classifier model features are generated from individual human subject data that includes: (a) a respective set of biomarker data and a respective set of clinical data; and (b) a COPD status of said individual human patient;wherein said human patient features match said random forest classifier model features; andgenerating, by the random forest classifier model, an output for said human patient indicating a COPD status of said patient, wherein the output indicates a likelihood of one composite COPD endpoint out of four possible composite COPD endpoints; and further comprising providing a recommendation for initiating treatment directed to a particular composite COPD endpoint.
2. The method of claim 1, wherein said composite COPD endpoint is chosen from a group comprising: (i) long-term oxygen utilization, (ii) long-term Non-Invasive Positive Pressure Ventilation (NIPPV), (iii) tracheostomy, or (iv) long-term nursing home placement.
3. The method of claim 1, wherein said receiving further comprises a respective set of demographic data for said human patient, and wherein said one or more features for input to the random forest classifier model is generated from said demographic data.
4. The method of claim 3, wherein said random forest classifier model determines whether there is a relationship between: (a) a plurality of features derived from at least: the respective set of biomarker data, the respective set of clinical data, and the respective set of demographic data; and (b) the COPD status of said patient.
5. The method of claim 1, wherein said biomarker data indicates a level of at least: receptor for advanced glycation end products (RAGE) and of club cell secretory- protein- 16 (CC-16).
6. The method of claim 1, wherein said biomarker data further includes a level of at least one of: interleukin-8 (IL-8), pulmonary surfactant protein-D (SP-D), C-reactive protein (CRP), eotaxin, and Alpha-1 antitrypsin (Al AT).
7. The method of claim 1, wherein said clinical data is selected from a group comprising: Global Initiative for Chronic Obstructive Lung Disease E (GOLD-E) phenoty pe over a prior year, Body Mass Index (BMI), age, height, long-term oxygen utilization, long-term Non-Invasive Positive Pressure Ventilation (NIPPV), tracheostomy, and long-term nursing home placement.
8. The method of claim 3, wherein said demographic data is selected from a group comprising: Body Mass Index (BMI), height, and age.
9. The method of claim 1, wherein said random forest classifier model has been trained using hyperparameter optimization using grid-search and 10-fold cross-validation of training data, and wherein hyperparameters include number of trees, tree depth, minimum samples per split and leaf, impurity criterion, and number of features considered.
10. The method of claim 9, further comprising sequentially removing clinical variables until loss of cross-validation performance on a training set is observed.
11. A system comprising:one or more processors; andone or more non-transitory computer-readable media, coupled to the one or more processors and storing instructions which, when executed by the one or more processors, cause the one or more processors to perform or control performance of operations comprising:receiving, for a human patient that has chronic obstructive pulmonary disease (COPD):a respective set of clinical data for said human patient; anda respective set of biomarker data collected from a sample from said human patient; andgenerating, from the respective data sets, one or more human patient features for input to a random forest classifier model;wherein said random forest classifier model has been trained using features generated from individual human subject data collected from a population of human subjects that have COPD;wherein at least one of said random forest classifier model features is generated from FEV1 and wherein said remaining random forest classifier model features are generated from individual human subject data that includes: (a) a respective set of biomarker data and a respective set of clinical data; and (b) a COPD status of said individual human patient;wherein said human patient features match said random forest classifier model features; andgenerating, by the random forest classifier model, an output for said human patient indicating a COPD status of said patient, wherein the output indicates a likelihood of one composite COPD endpoint out of four possible composite COPD endpoints; and further comprising providing a recommendation for initiating treatment directed to a particular composite COPD endpoint.
12. The method of claim 11, wherein said composite COPD endpoint is chosen from a group comprising: (i) long-term oxygen utilization, (ii) long-term Non-Invasive Positive Pressure Ventilation (NIPPV), (iii) tracheostomy, or (iv) long-term nursing home placement.
13. The method of claim 11, wherein said receiving further comprises a respective set of demographic data for said human patient, and wherein said one or more features for input to the random forest classifier model is generated from said demographic data.
14. The method of claim 13, wherein said random forest classifier model determines whether there is a relationship between: (a) a plurality of features derived from at least: the respective set of biomarker data, the respective set of clinical data, and the respective set of demographic data; and (b) the COPD status of said patient.
15. The method of claim 11, wherein said biomarker data indicates a level of at least: receptor for advanced glycation end products (RAGE) and of club cell secretory protein-16 (CC-16).
16. The method of claim 11, wherein said biomarker data further includes a level of at least one of interleukin-8 (IL-8), pulmonary surfactant protein-D (SP-D). C-reactive protein (CRP), eotaxin, and Alpha- 1 antitrypsin (Al AT).
17. The method of claim 11, wherein said clinical data is selected from a group comprising: Global Initiative for Chronic Obstructive Lung Disease E (GOLD-E) phenotype over a prior year, Body Mass Index (BMI), age, height, long-term oxygen utilization, long-term Non-Invasive Positive Pressure Ventilation (NIPPV), tracheostomy, and long-term nursing home placement.
18. The method of claim 13, wherein said demographic data is selected from a group comprising: Body Mass Index (BMI), height, and age.
19. The method of claim 11, wherein said random forest classifier model has been trained using hyperparameter optimization using grid-search and 10-fold cross-validation of training data, and wherein hyperparameters include number of trees, tree depth, minimum samples per split and leaf, impurity criterion, and number of features considered.
20. The method of claim 19, further comprising sequentially removing clinical variables until loss of cross-validation performance on a training set is observed.