Method for constructing a screening model for chronic obstructive pulmonary disease based on quantitative CT

By constructing a multimodal screening model based on quantitative CT, the problem of early diagnosis of COPD has been solved, efficient and accurate screening and diagnosis have been achieved, the management level of COPD has been improved, and it is suitable for different populations and environments, providing personalized treatment plans.

CN119230131BActive Publication Date: 2025-09-12THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411234750.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-09-12
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

With existing technologies, early diagnosis of chronic obstructive pulmonary disease (COPD) is difficult to popularize in primary healthcare settings. Traditional CT evaluation is subjective and labor-intensive, and existing screening tools have limited accuracy and cannot effectively identify high-risk groups.

Method used

A quantitative CT-based COPD screening model was constructed. By screening a derivative cohort and an external validation cohort within the study population, single-modality and multimodal modeling were performed by combining questionnaires, pulmonary function tests, and chest CT scans. The model was optimized using the XGBoost algorithm, and the questionnaire data, quantitative CT data, and structured sign data were integrated for model validation and evaluation.

Benefits of technology

It has significantly improved the early screening and diagnosis of COPD, improved patients' quality of life and prognosis. The model performs well in different regions and ethnic groups, and its recognition ability is better than traditional tools. It provides personalized management plans and reduces unnecessary examinations and medical expenses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119230131B_ABST
    Figure CN119230131B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing a quantitative CT-based chronic obstructive pulmonary disease screening model, comprising: screening a study population according to preset criteria to obtain a derivation cohort and a first external validation cohort; conducting questionnaires, pulmonary function tests, and chest CT scan analysis on subjects in the derivation cohort to obtain a training dataset and an internal validation dataset; collecting clinical data, performing pulmonary function tests, and chest CT scans on subjects in the first external validation cohort, and combining the data with a second external validation cohort to obtain an external validation dataset; performing single-modality modeling and multimodality modeling based on the types of data in the training dataset, and validating and evaluating the modeled models using the internal validation dataset to obtain the optimal model; and performing performance evaluation of the optimal model using the external validation dataset. The present invention can significantly improve the early screening, diagnosis, and management of chronic obstructive pulmonary disease, thereby improving the quality of life and prognosis of patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disease screening, and in particular to a method for constructing a chronic obstructive pulmonary disease screening model based on quantitative CT. Background Art

[0002] Underdiagnosis of chronic obstructive pulmonary disease (COPD) is a global problem, particularly in developing countries. According to a large epidemiological study in my country, approximately 60.2% of COPD patients did not self-report typical symptoms, and only 12% of COPD patients had undergone pulmonary function testing (PFT) prior to the survey, demonstrating the hidden nature of COPD's early symptoms and low public awareness of its presence. However, despite being the current "gold standard" for COPD diagnosis, PFTs are not widely used in primary care settings. Common reasons include a lack of standardized PFT techniques, the high cost and time required for PFTs, and limited ability to interpret results. Early diagnosis of COPD is essential for the timely initiation of appropriate lifestyle and treatment interventions to reduce morbidity, delay disease progression, and improve patients' quality of life. Therefore, alternative strategies are urgently needed to provide rapid detection and accurate assessment of COPD for optimal clinical decision-making.

[0003] It is well known that systematic active case finding in primary care settings using mailed screening questionnaires is an effective approach. A widely used tool, endorsed by Chinese clinical diagnosis and treatment guidelines, is the COPD Screening Questionnaire (COPD-SQ). Other tools include the COPD Population Screening Questionnaire (COPD-PS) and the COPD Diagnostic Questionnaire (CDQ). These simple, cost-effective methods for screening for COPD based on symptoms and exposure factors are crucial for identifying high-risk individuals. However, the accuracy of questionnaires is difficult to significantly improve. Some even include multiple subjective questions, which can lead to subject recall bias and are time-consuming and labor-intensive to collect. The generalizability of different questionnaires requires further validation. In contrast, chest computed tomography (CT) is widely available, particularly in large-scale lung cancer screening. Furthermore, the Global Initiative for Chronic Obstructive Lung Disease (GOLD) 2024 emphasizes the importance of CT in the assessment of patients with stable COPD and the use of lung cancer imaging for COPD screening, highlighting the role of imaging. This evidence-based recommendation is based on the limitations of PFTs. CT, with its advantages of being noninvasive and high-resolution, can be used to assess changes in the lung parenchyma, airways, and pulmonary vasculature that occur with declining lung function and aging. However, the traditional manual reading method has certain limitations when using CT images to evaluate COPD. On the one hand, the interpretation of CT is somewhat subjective and will be affected by factors such as different judgment standards and uneven reading levels of radiologists. It requires long-term systematic training to correct and reduce assessment bias. On the other hand, COPD is a highly heterogeneous disease that cannot be finely analyzed through single-dimensional images. All-round, multi-dimensional visual observation is bound to increase the workload of medical staff and may even cause problems due to different CT imaging effects on different scanning devices due to differences in detection parameters, reconstruction algorithms, image color differences, etc. Therefore, with the help of rapidly developing artificial intelligence (AI) technology, it is crucial to find an objective, efficient, and accurate intelligent COPD assessment method for CT whole-lung level.

[0004] Radiomics is a relatively new method that can rapidly extract quantitative features from medical images (such as CT) and has shown great potential for clinical decision-making. Early studies performed segmentation and quantification by manually marking regions of interest (typically lesion sites). However, manual segmentation of diffuse and heterogeneous lung lesions (such as emphysema, bronchiolitis, and interstitial lung disease) is difficult and time-consuming due to unclear boundaries or low contrast on CT images. The error rate can also be high due to operator inexperience. Therefore, automated segmentation of the entire lung region will help comprehensively quantify lung abnormalities in COPD patients and assist in clinical treatment decisions.

[0005] Of note, two recent studies have suggested that whole-lung CT radiomics indices are suitable for intelligent COPD screening. However, one study used a smoking-related Western cohort to train the model, which may not be applicable to other regions. While the other developed a model using a Chinese population, the cohort consisted primarily of patients with established COPD, who exhibit significant differences in clinical phenotype, biochemical markers, and CT imaging compared to the general population. Furthermore, these indices may be affected by imaging processing algorithms and performance, making them difficult to apply in routine practice. Therefore, there is a need for more appropriate CT assessment protocols for COPD screening in my country. Summary of the Invention

[0006] In order to overcome the deficiencies of the prior art, the present invention aims to provide a method for constructing a COPD screening model based on quantitative CT.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A method for constructing a COPD screening model based on quantitative CT, comprising:

[0009] The study population was screened according to pre-specified criteria to obtain the derivation cohort and the first external validation cohort;

[0010] Conducting questionnaires, pulmonary function tests, and chest CT scans on the subjects in the derivation cohort to obtain a training dataset and an internal validation dataset; conducting clinical data collection, pulmonary function tests, and chest CT scans on the subjects in the first external validation cohort, and combining the data with the pre-set second external validation cohort to obtain an external validation dataset;

[0011] Performing single-modal modeling and multi-modal modeling according to the type of each data in the training data set, and verifying and evaluating the modeled model using the internal validation data set to obtain the optimal model;

[0012] The performance of the optimal model is evaluated using the external validation dataset.

[0013] Preferably, the optimal model is constructed based on the XGBoost algorithm.

[0014] Preferably, screening is performed in the study population according to preset criteria to obtain a derivation cohort and a first external validation cohort, including:

[0015] The subjects from the derivation cohort who met the pre-set inclusion and exclusion criteria were randomly divided into a training cohort and an internal validation cohort at a ratio of 8:2;

[0016] The preset hospital cohorts were used as the first independent external validation cohorts.

[0017] Preferably, the subjects of the derivation cohort are subjected to questionnaires, pulmonary function tests, and chest CT scans to obtain a training dataset and an internal validation dataset, and the subjects of the first external validation cohort are subjected to clinical data collection, pulmonary function tests, and chest CT scans, and combined with the preset second external validation cohort to obtain an external validation dataset, including:

[0018] A questionnaire is administered to the subjects of the derivation cohort to obtain questionnaire data; the questionnaire indicators include: demographic characteristics, anthropometric information, smoking history, personal and family medical history, respiratory symptoms, and lifestyle and dietary patterns;

[0019] The gender, age, smoking history, pulmonary function test data, and peripheral blood eosinophil count of all subjects in the first external validation cohort were collected through the electronic medical record system to obtain clinical data;

[0020] The data of the pre-specified second external validation cohort were obtained through public databases; the data of the second external validation cohort included demographic characteristics, smoking history, medical history, and chest CT images;

[0021] performing pulmonary function tests and chest CT scans on the subjects of the derivation cohort and the subjects of the first external validation cohort, respectively, to obtain pulmonary function test data and chest CT images;

[0022] performing lung parenchyma analysis and airway analysis on the chest CT image using an image registration algorithm to obtain quantitative CT data;

[0023] Performing structured processing on the diagnostic results according to the CT report corresponding to the chest CT image to obtain structured sign data; the structured sign data includes: whether there are bullae, chronic inflammation, fibrous foci, calcified foci, nodules, bronchiectasis, tuberculosis, bronchitis, and pleural thickening;

[0024] Determining the questionnaire data, the pulmonary function test data, the quantitative CT data, and the structured sign data corresponding to the subjects of the derivation cohort as the training data set and the internal validation data set;

[0025] The clinical data, the pulmonary function test data, the quantitative CT data, the structured sign data corresponding to the subjects of the first external validation cohort and the data of the second external validation cohort are determined as the external validation data set.

[0026] Preferably, the single-modal modeling includes: modeling using the questionnaire data in the training data set, modeling using the quantitative CT data in the training data set, and modeling using the structured sign data in the training data set; the multi-modal modeling includes: modeling using the questionnaire data and quantitative CT data in the training data set, modeling using the questionnaire data and structured sign data in the training data set, modeling using the quantitative CT data and structured sign data in the training data set, and modeling using the questionnaire data, quantitative CT data, and structured signs in the training data set.

[0027] Preferably, the indicators for validating and evaluating the model after modeling using the internal validation dataset include: AUC, sensitivity, specificity, accuracy, positive predictive value, negative predictive value and F1 score.

[0028] Preferably, the quantitative CT data includes: LAA%-950, LAA%-950 left upper lung, LAA%-950 left lower lung, LAA%-950 right upper lung, LAA%-950 right middle lung, LAA%-950 right lower lung, LAA%-910 left lower lung, LAA%-910 right lower lung, maximum diameter of the first-level airway lumen and average diameter of the fourth-level airway lumen.

[0029] Preferably, the questionnaire data, the pulmonary function test data, the quantitative CT data, and the structured sign data corresponding to the subjects of the derivation cohort are determined as the training dataset and the internal validation dataset, further comprising:

[0030] The dataset consisting of the questionnaire data, quantitative CT data, and structured sign data corresponding to the subjects of the derivation cohort was randomly divided into a training cohort and an internal validation cohort in a ratio of 80% and 20% using the createDataPartition function of the caret R package;

[0031] Dividing the data of the training cohort and the internal validation cohort into measurement data and counting data, respectively, and further dividing the counting data into binary data and ordered multi-classification data;

[0032] Determine the measurement data as the first sequence;

[0033] The values ​​of each binary count data are recoded as 1 and 2 to obtain the second series;

[0034] Re-encoding the ordered multi-classification data by adopting a fusion weight encoding method to obtain a third sequence;

[0035] Use the missforest R package to interpolate missing values ​​for the first sequence, the second sequence, and the third sequence in the training cohort and the internal validation cohort, respectively, to obtain interpolated data sequences;

[0036] Using the nearZeroVar function of the caret R package, identifying and removing features of zero variance and near-zero variance in the interpolated data sequence corresponding to the training cohort, to obtain a data sequence after removal;

[0037] The createFolds function is used to randomly split the eliminated data sequence corresponding to the training queue into multiple groups, and finally a preprocessed training data set and a preprocessed internal validation data set are obtained.

[0038] Preferably, the formula for the fusion weight encoding is: ;in, is the maximum integer value after recoding, is the total number of categories, For category For the target variable The contribution weight of ;in, is an event of category i, is the probability that category i and target variable Y occur simultaneously, and They are and the marginal probability of Y, is the value of category i, and y is the specific value of the target variable Y.

[0039] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0040] The present invention provides a method for constructing a quantitative CT-based chronic obstructive pulmonary disease screening model, comprising: screening a study population according to preset criteria to obtain a derivative cohort and a first external validation cohort; conducting questionnaires, pulmonary function tests, and chest CT scan analysis on subjects in the derivative cohort to obtain a training data set and an internal validation data set; collecting clinical data, performing pulmonary function tests, and chest CT scans on subjects in the first external validation cohort, and combining the data with a preset second external validation cohort to obtain an external validation data set; performing single-modality modeling and multi-modality modeling according to the type of each data in the training data set, and validating and evaluating the modeled model using the internal validation data set to obtain an optimal model; and performing performance evaluation on the optimal model using the external validation data set. The present invention can significantly improve the level of early screening, diagnosis, and management of chronic obstructive pulmonary disease, thereby improving the quality of life and prognosis of patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 A flowchart of a method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0044] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0045] Figure 1 A flow chart of the method provided in the embodiment of the present invention is shown in FIG. Figure 1 As shown, the present invention provides a method for constructing a COPD screening model based on quantitative CT, comprising:

[0046] Step 1: Screen the study population according to the preset criteria to obtain the derivation cohort and the first external validation cohort;

[0047] Step 2: Conduct questionnaire surveys, pulmonary function tests, and chest CT scans on the subjects in the derivation cohort to obtain a training dataset and an internal validation dataset. Furthermore, conduct clinical data collection, pulmonary function tests, and chest CT scans on the subjects in the first external validation cohort, and combine the data with the pre-set second external validation cohort to obtain an external validation dataset.

[0048] Step 3: Performing unimodal modeling and multimodal modeling according to the type of each data in the training dataset, and using the internal validation dataset to verify and evaluate the modeled model to obtain the optimal model;

[0049] Step 4: Use the external validation dataset to evaluate the performance of the optimal model.

[0050] Preferably, the optimal model is constructed based on the XGBoost algorithm.

[0051] Preferably, the study population is screened according to preset criteria to obtain a derivation cohort and an external validation cohort, including:

[0052] The subjects from the derivation cohort who met the pre-set inclusion and exclusion criteria were randomly divided into a training cohort and an internal validation cohort at a ratio of 8:2;

[0053] The preset hospital cohorts were used as the first independent external validation cohorts.

[0054] Specifically, all participants from China were retrospectively recruited between April 2017 and May 2024 from four centers: four neighborhoods in Guangzhou, the First Affiliated Hospital of Guangzhou Medical University, Xiangyang Central Hospital, and the Second Affiliated Hospital of Xi'an Jiaotong University. Inclusion criteria were as follows: 1) age 35-80 years; 2) completion of at least one PFT; 3) inspiratory chest CT scan performed in the supine position; 4) chest CT scan and PFT performed at the same center; and 5) community-based participants completed a respiratory epidemiology questionnaire. Exclusion criteria were as follows: 1) incomplete CT images; 2) extensive imaging artifacts; 3) history of lung resection; and 4) questionnaire information missing >50%. A total of 3,653 participants were enrolled. Of these, 1,950 community-based participants (650 COPD patients and 1,300 controls, fully frequency-matched by sex and age) were assigned to the derivation cohort and randomly divided in an 8:2 ratio into a training cohort (n = 1,560) and an internal validation cohort (n = 390). 1186 patients from the First Affiliated Hospital of Guangzhou Medical University, 225 patients from Xiangyang Central Hospital, and 292 patients from the Second Affiliated Hospital of Xi'an Jiaotong University were assigned to external validation cohorts 1, 2, and 3, respectively. A randomized subset of the US NLST cohort (n = 453) was further used as external validation cohort 4 to verify the multiethnic extrapolation and robustness of low-dose CT scanning. This study was approved by all institutional review boards (approval numbers: ES–2024–K060–01; NLST–1167). Due to the retrospective nature of the study, written informed consent was not required.

[0055] The derivation cohort used a forced expiratory volume in one second (FEV1) / forced vital capacity (FVC) ratio <0.7 before inhaled bronchodilators as the diagnostic criterion for COPD. The external validation cohort (excluding the NLST cohort) used a FEV1 / FVC ratio <0.7 after inhaled bronchodilators as the diagnostic criterion for COPD. Further stratification was based on the GOLD guidelines: GOLD stage 1, FEV1%pred ≥ 80%; GOLD stage 2: 50% ≤ FEV1%pred < 80%; GOLD stage 3: 30% ≤ FEV1%pred < 50%; and GOLD stage 4: FEV1%pred < 30%.

[0056] Furthermore, this embodiment performs sample size calculation, which is as follows:

[0057] This trial employed a two-gate diagnostic design with qualitative outcome measures. The sample size for the case and control groups was equal. Key parameters were as follows: sensitivity = 80%, with a tolerance for sensitivity of 5%; specificity = 70%, with a tolerance for specificity of 5%. A confidence level of 1-α = 0.95 was selected, and the prevalence ratio (a priori probability) in this trial was 0.33 (COPD group:control group = 1:2). The null sensitivity and specificity were both 0.5.

[0058] According to PASS software calculations, a total of 800 samples were required, including 264 patients whose COPD diagnosis was clearly excluded by the gold standard, to achieve the predefined 95% confidence interval for the test. Furthermore, accounting for a 20% dropout rate, a final sample size of 1000 was required. A total of 1950 subjects were enrolled in this experimental model development, with 650 in the COPD group and 1300 in the control group, meeting the minimum sample size requirement for model training.

[0059] Preferably, the subjects of the derivation cohort are subjected to questionnaires, pulmonary function tests, and chest CT scans to obtain a training dataset, and the subjects of the external validation cohort are subjected to clinical data collection, pulmonary function tests, and chest CT scans to obtain an external validation dataset, including:

[0060] A questionnaire is administered to the subjects of the derivation cohort to obtain questionnaire data; the questionnaire indicators include: demographic characteristics, anthropometric information, smoking history, personal and family medical history, respiratory symptoms, and lifestyle and dietary patterns;

[0061] The gender, age, smoking history, pulmonary function test data, and peripheral blood eosinophil count of all subjects in the first external validation cohort were collected through the electronic medical record system to obtain clinical data;

[0062] The data of the pre-specified second external validation cohort were obtained through public databases; the data of the second external validation cohort included demographic characteristics, smoking history, medical history, and chest CT images;

[0063] Pulmonary function tests and chest CT scans were performed on the subjects of the derivation cohort and the subjects of the first external validation cohort, respectively, to obtain pulmonary function test data and chest CT images; the pulmonary function test data was used to assign disease labels to the subjects and was not included in the modeling process.

[0064] performing lung parenchyma analysis and airway analysis on the chest CT image using an image registration algorithm to obtain quantitative CT data;

[0065] Performing structured processing on the diagnostic results according to the CT report corresponding to the chest CT image to obtain structured sign data; the structured sign data includes: whether there are bullae, chronic inflammation, fibrous foci, calcified foci, nodules, bronchiectasis, tuberculosis, bronchitis, and pleural thickening;

[0066] Determining the questionnaire data, the pulmonary function test data, the quantitative CT data, and the structured sign data corresponding to the subjects of the derivation cohort as the training data set and the internal validation data set;

[0067] The clinical data, the pulmonary function test data, the quantitative CT data, the structured sign data corresponding to the subjects of the first external validation cohort and the data of the second external validation cohort are determined as the external validation data set.

[0068] Specifically, the steps of data collection and processing in this embodiment are as follows:

[0069] 1) Questionnaire and clinical data collection

[0070] Participants in the derivation cohort completed an offline questionnaire under the guidance of staff experienced in epidemiological survey techniques. The questionnaire contained 40 items, covering demographic characteristics, anthropometric information, smoking history, personal and family medical history, respiratory symptoms, and lifestyle and dietary patterns. For the external validation cohort, sex, age, smoking history, pulmonary function test data, and peripheral blood eosinophil counts were collected from all participants from China through the electronic medical record system. Data from the US NLST external validation cohort, including demographic characteristics, smoking history, medical history, and chest CT images, were provided through the US NIH website.

[0071] 2) Pulmonary function test

[0072] All subjects underwent pulmonary function testing using a domestically produced, new differential pressure spirometer, the Youxi PF680 (Yiliankang, Zhejiang, China) or a MasterScreen Pneumo spirometer (Jaeger, Bavaria, Germany). Participants in the external validation cohort from China underwent bronchodilator testing. All pulmonary function tests were performed by professional technicians. The primary measures of FEV1 / FVC and FEV1%pred were included in this study.

[0073] 3) Chest CT Scan Protocol

[0074] All cohorts used CT scanners from seven manufacturers: SIEMENS, GE, NMS, TOSHIBA, Philips, UIH, and Anke. Chest CT scans were performed at the end of maximal inspiration with the patient in the supine position with arms raised overhead. The scan area extended from the thoracic inlet to the level of the bilateral adrenal glands. All images were exported in Digital Imaging and Communications in Medicine (DICOM) format for quantitative data extraction.

[0075] 4) Image analysis

[0076] CT images were analyzed using NeuLungCare–QA software (v1.0, Neusoft Software Co., Ltd., Shenyang), which was commercially certified and validated in a previous study (13) for quantification of CT emphysema and airway lesions. A total of 47 QCT features were extracted in this study. Briefly, CT emphysema was quantified for the entire lung and for each lung lobe using low attenuation areas below −950 (LAA-950) and −910 (LAA-910) Hounsfield units on full inspiratory images, respectively.

[0077] CT airway measurements were generated from bronchi from generation 0 to generation 4. Percent airway wall area (%WA) was measured by dividing the total bronchial cross-sectional area by the wall area. For airway size measurements, airway length was defined as the distance from the start to the end of the branch point by placing a smooth centerline through the lumen. Wall thickness (WT) and lumen diameter (LD) were calculated for each airway every 2 mm. Maximum, minimum, and mean values ​​for WT and LD were generated from five measurements.

[0078] All original chest CT images in DICOM format were imported into the NeuLungCARE-QA system (version 1.0, Neusoft Software Co., Ltd., Shenyang) for quantification of CT emphysema and airway lesions. The analysis process is briefly described as follows: the system processes the input CT images using the optimized Demons registration algorithm and performs quantitative measurements according to the classification criteria of previous literature:

[0079] (1) Lung parenchyma analysis: Each lung lobe was automatically segmented along the interlobar fissures, with -950 Hounsfield units (HU) and -910 HU set as thresholds, respectively. Lung attenuation voxels with a value less than -950 HU or -910 HU were considered low-attenuation areas, expressed as LAA%-950 and LAA%-910, and the LAA% of the whole lung, left upper lobe, left lower lobe, right upper lobe, right middle lobe, and right lower lobe were calculated respectively.

[0080] (2) Airway analysis: The bronchial tree was reconstructed in three dimensions, and then the airways of levels 0-4 were automatically selected as target airways. The percentage of airway wall area (%WA) was obtained by dividing the total bronchial cross-sectional area by the wall area. For airway size measurement, airway length was defined as the distance from the start point to the end point of the branch point by placing a smooth center line in the lumen. The wall thickness (WT) and lumen diameter (LD) of each airway were calculated every 2 mm. The maximum, minimum, and average values ​​of WT and LD were generated from five measurements.

[0081] 5) CT report text processing

[0082] This embodiment simultaneously collects CT reports interpreted by professional radiologists, structures the diagnostic results, and ultimately records nine signs based on the presence or absence of pulmonary bullae (including emphysema), chronic inflammation (including acute and chronic infection), fibrous foci, calcification foci, nodules (including granulomas, neoplasms, and suspected tumors), bronchiectasis, tuberculosis, bronchitis (including bronchiolitis and chronic inflammation of the small airways), and pleural thickening.

[0083] Preferably, before performing single-modal modeling and multi-modal modeling, the method further includes:

[0084] The dataset consisting of the questionnaire data, quantitative CT data, and structured sign data corresponding to the subjects of the derivation cohort was randomly divided into a training cohort and an internal validation cohort in a ratio of 80% and 20% using the createDataPartition function of the caret R package;

[0085] Dividing the data of the training cohort and the internal validation cohort into measurement data and counting data, respectively, and further dividing the counting data into binary data and ordered multi-classification data;

[0086] Determine the measurement data as the first sequence;

[0087] The values ​​of each binary count data are recoded as 1 and 2 to obtain the second series;

[0088] Re-encoding the ordered multi-classification data by adopting a fusion weight encoding method to obtain a third sequence;

[0089] Use the missforest R package to interpolate missing values ​​for the first sequence, the second sequence, and the third sequence in the training cohort and the internal validation cohort, respectively, to obtain interpolated data sequences;

[0090] Using the nearZeroVar function of the caret R package, identifying and removing features of zero variance and near-zero variance in the interpolated data sequence corresponding to the training cohort, to obtain a data sequence after removal;

[0091] The createFolds function is used to randomly split the eliminated data sequence corresponding to the training queue into multiple groups, and finally a preprocessed training data set and a preprocessed internal validation data set are obtained.

[0092] Preferably, the formula for the fusion weight encoding is: ;in, is the maximum integer value after recoding, is the total number of categories, For category For the target variable The contribution weight of ;in, is an event of category i, is the probability that category i and target variable Y occur simultaneously, and They are and the marginal probability of Y, is the value of category i, and y is the specific value of the target variable Y.

[0093] Furthermore, in this embodiment, for binary count data, in addition to simply encoding it as 1 and 2, a robust encoding method for noise and outliers can be introduced. For example, quantiles can be calculated to detect and redistribute outliers, ensuring more robust encoding.

[0094] For ordered multi-class data, weighted encoding is introduced to encode the classification values ​​into continuous positive integers based on the contribution of each classification in the prediction of the target variable (which can be determined by preliminary feature importance analysis).

[0095] In addition, this embodiment can also adopt an optimized data segmentation strategy, namely Adaptive Stratified Partitioning, which not only relies on a simple 80 / 20 split but also achieves a more balanced split in the distribution of target variables and key features. Indicators such as weighted Kappa are used to ensure data distribution consistency in the training and validation sets.

[0096] Furthermore, this embodiment introduces Neighborhood Dynamic Regression Imputation (NDRI), an enhanced version of the classic missForest method. This method combines local linear regression with the K-nearest neighbor algorithm to achieve dynamic prediction of missing values.

[0097] In addition, this example uses an Augmented Generative Adversarial Network (AGAN) for data augmentation to generate diverse samples within the feature space, thereby improving the model's generalization capabilities. Furthermore, multi-level cross-validation is implemented. Beyond the classic ten-fold approach, multi-scale partitioning and subsampling strategies are introduced for multi-level feature and model validation to reduce the risk of overfitting.

[0098] Preferably, the single-modality modeling includes: modeling using questionnaire data in the training data set, modeling using quantitative CT data in the training data set, and modeling using CT reports in the training data set; the multi-modality modeling includes: modeling using questionnaire data and quantitative CT data in the training data set, modeling using questionnaire data and structured sign data in the training data set, modeling using quantitative CT data and structured sign data in the training data set, and modeling using questionnaire data, quantitative CT data, and structured signs in the training data set.

[0099] Preferably, the model is verified and evaluated to obtain the optimal model, including:

[0100] The internal validation cohort is used to validate and evaluate the model after modeling to obtain the optimal model; the evaluation indicators include: AUC, sensitivity, specificity, accuracy, positive predictive value, negative predictive value and F1 score.

[0101] Furthermore, this embodiment also includes:

[0102] 7) Model construction and performance evaluation

[0103] The XGBoost algorithm was used as the core of the screening model. The specific construction method was as follows: Based on the ten-fold cross-validation framework, the BayesianOptimization function of the rBayesianOptimization R package was used to automatically optimize the model hyperparameters, including eta, max_depth, subsample, colsample_byTree, gamma, min_child_weight, and nrounds, on the preprocessed training cohort. The parameters init_points and n_iter were set to 5 and 1, respectively, and the remaining parameters were set to their default values. AUC was used as the cross-validation loss function value for evaluation. After determining the optimal hyperparameters, SHAP analysis was used to screen out the top ten features, and the final screening model was rebuilt in the same way.

[0104] To identify the optimal screening strategy, the aforementioned workflow was used to model single-modality (questionnaire, quantitative CT, and CT report) and multimodality (questionnaire + quantitative CT, questionnaire + CT report, quantitative CT + CT report, and questionnaire + quantitative CT + CT report) strategies. Finally, these strategies were evaluated on an internal validation cohort, and the optimal model was designated AutoCOPD. The area under the curve (AUC) was used as an overall discriminant metric for model classification performance. Additionally, sensitivity, specificity, accuracy, positive predictive value, negative predictive value, and F1 score were used as additional evaluation metrics. Finally, the area under the receiver operating characteristic curve (AUC) of the different models was compared using the DeLong test; a P < 0.05 indicated statistically significant differences in AUC. These metrics were calculated using the pROCR package. The performance of AutoCOPD was further evaluated on all external validation cohorts. Lowess curves were plotted using the stats R package to reflect model calibration. The Hosmer-Lemeshow test was performed using the ResourceSelection R package; a P > 0.05 indicated good model calibration. The Brier score was then calculated to comprehensively assess calibration. Clinical decision curves (DCAs) were constructed for the model across all cohorts using the rmda R package. Subgroup analyses were performed in the internal and external validation cohorts to clarify the applicability of the model. The COPD identification performance of the model was compared with that of the COPD-SQ, a classic screening tool, in the internal validation cohort.

[0105] 8) Core Results

[0106] Ultimately, this example identified an XGBoost model based on 10 quantitative CT features as the optimal screening model (i.e., AutoCOPD). These features included LAA%-950, LAA%-950 left upper lung, LAA%-950 left lower lung, LAA%-950 right upper lung, LAA%-950 right middle lung, LAA%-950 right lower lung, LAA%-910 left lower lung, LAA%-910 right lower lung, maximum diameter of the first-stage airway lumen, and mean diameter of the fourth-stage airway lumen. This model was well validated across various scenarios, regions, and ethnic populations, with good model calibration. Decision curve analysis showed an overall clinical net benefit range of 0.12-0.66, and overall recognition capability significantly outperformed the COPD-SQ. In summary, AutoCOPD can serve as a novel COPD screening tool.

[0107] 8) System Development

[0108] This example also uses R Shiny to create a user-friendly, free online app based on the AutoCOPD model. This app includes routine, custom, and individual analyses. In routine analysis, users can create their own datasets using the sample file format and upload them to the app for batch prediction. A prediction score ≥50% indicates a COPD patient, while a score <50% indicates a non-COPD patient. Custom analysis is designed for external validation of self-created datasets, providing an optimal decision threshold based on the Youden index and displaying standard evaluation metrics, receiver operating characteristic (ROC) curves, calibration curves, and decision curves. All missing values ​​are automatically imputed before analysis, and users can download the imputed data and prediction results for local research. In individual prediction, upon entering the actual values ​​for the 10 features required by the model (at least one feature value must be entered) and clicking the "Calculate" button, the app automatically predicts an individual's risk of COPD. A SHAP force plot is also displayed to indicate feature contribution: purple features push the prediction toward "non-COPD," while yellow features push the prediction toward "COPD." Overall, the app is easy to use, responsive, and focused, making it highly compatible with clinical practice.

[0109] The advantages of the quantitative CT-based chronic obstructive pulmonary disease (COPD) screening model of the present invention are mainly reflected in the following aspects:

[0110] (1) Early screening and diagnosis: Through a machine learning model that only includes 10 whole-lung inspiratory phase quantitative CT indicators, COPD can be detected early, thereby increasing the possibility of early intervention and treatment and slowing down disease progression.

[0111] (2) Multimodal comparison, integrating different data sources (such as questionnaires, quantitative CT, and structured features) to make a more comprehensive comparison and derive the optimal screening plan suitable for my country.

[0112] (3) Quantitative analysis: Commercialized quantitative CT analysis software certified by the China Food and Drug Administration can accurately measure changes in lung structure, such as emphysema, airway structure and volume changes, providing objective data support and improving the credibility of screening.

[0113] (4) Standardized process: preset standard patient screening and data collection process ensures the homogeneity of the study population and the comparability of the data, and improves the reliability and validity of the results.

[0114] (5) External validation: By evaluating the external validation cohort, the generalization ability of the model can be effectively tested to ensure the effectiveness of the model in different populations and environmental conditions.

[0115] (6) Potential for personalized medicine. This model has strong clinical interpretability, which is conducive to combining other individual characteristics and clinical data during clinical diagnosis and treatment to tailor personalized management plans for subjects.

[0116] (7) Promote clinical decision-making. The construction and evaluation of the model will provide a scientific basis for clinicians, helping them make more accurate diagnosis and treatment decisions, thereby improving treatment outcomes.

[0117] (8) Resource optimization: accurate screening models can help clinical institutions optimize resource allocation, improve screening efficiency for high-risk populations, and reduce unnecessary examinations and medical expenses.

[0118] (9) Promote research progress. The construction of the model of the present invention provides a foundation for the study of COPD, helps to understand the disease progression mechanism and risk factors, and provides data support and theoretical basis for future related research.

[0119] Through the above advantages, the model constructed by the present invention can significantly improve the early screening, diagnosis and management of chronic obstructive pulmonary disease, thereby improving the quality of life and prognosis of patients.

[0120] The beneficial effects of the present invention are as follows:

[0121] (1) The study involved in this invention is the largest multicenter quantitative CT-COPD screening model study in China and abroad;

[0122] (2) Using 10 highly interpretable whole-lung inspiratory phase chest quantitative CT features, we successfully constructed an intelligent COPD screening model (AutoCOPD) and online app based on the XGBoost algorithm. The screening performance in the training cohort, internal validation cohort, and external validation cohort was good, and the recognition performance in a highly heterogeneous community population was better than that of the classic screening tool COPD-SQ.

[0123] (3) The model of the present invention has good COPD identification capabilities before and after receiving bronchodilators, in different regions and among different ethnic groups, laying a new theoretical foundation for large-scale COPD imaging screening.

[0124] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0125] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A method for constructing a chronic obstructive pulmonary disease screening model based on quantitative CT, characterized in that: include: The study population was screened according to pre-specified criteria to obtain the derivation cohort and the first external validation cohort; Conducting questionnaires, pulmonary function tests, and chest CT scans on the subjects in the derivation cohort to obtain a training dataset and an internal validation dataset; conducting clinical data collection, pulmonary function tests, and chest CT scans on the subjects in the first external validation cohort, and combining the data with the pre-set second external validation cohort to obtain an external validation dataset; Performing single-modal modeling and multi-modal modeling according to the type of each data in the training data set, and verifying and evaluating the modeled model using the internal validation data set to obtain the optimal model; Performing performance evaluation on the optimal model using the external validation dataset; The subjects of the derivation cohort were subjected to questionnaires, pulmonary function tests, and chest CT scans to obtain a training dataset and an internal validation dataset. Clinical data collection, pulmonary function tests, and chest CT scans were performed on the subjects of the first external validation cohort, and the external validation dataset was obtained by combining the data from the preset second external validation cohort, including: A questionnaire is administered to the subjects of the derivation cohort to obtain questionnaire data; the questionnaire indicators include: demographic characteristics, anthropometric information, smoking history, personal and family medical history, respiratory symptoms, and lifestyle and dietary patterns; The gender, age, smoking history, pulmonary function test data, and peripheral blood eosinophil count of all subjects in the first external validation cohort were collected through the electronic medical record system to obtain clinical data; The data of the pre-specified second external validation cohort were obtained through public databases; the data of the second external validation cohort included demographic characteristics, smoking history, medical history, and chest CT images; performing pulmonary function tests and chest CT scans on the subjects of the derivation cohort and the subjects of the first external validation cohort, respectively, to obtain pulmonary function test data and chest CT images; performing lung parenchyma analysis and airway analysis on the chest CT image using an image registration algorithm to obtain quantitative CT data; Performing structured processing on the diagnostic results according to the CT report corresponding to the chest CT image to obtain structured sign data; the structured sign data includes: whether there are bullae, chronic inflammation, fibrous foci, calcified foci, nodules, bronchiectasis, tuberculosis, bronchitis, and pleural thickening; Determining the questionnaire data, the pulmonary function test data, the quantitative CT data, and the structured sign data corresponding to the subjects of the derivation cohort as the training data set and the internal validation data set; Determining the clinical data, the pulmonary function test data, the quantitative CT data, the structured sign data corresponding to the subjects of the first external validation cohort and the data of the second external validation cohort as the external validation data set; The quantitative CT data include: LAA%-950, LAA%-950 left upper lung, LAA%-950 left lower lung, LAA%-950 right upper lung, LAA%-950 right middle lung, LAA%-950 right lower lung, LAA%-910 left lower lung, LAA%-910 right lower lung, maximum diameter of the first-stage airway lumen, and average diameter of the fourth-stage airway lumen; Determining the questionnaire data, the pulmonary function test data, the quantitative CT data, and the structured sign data corresponding to the subjects of the derivation cohort as the training dataset and the internal validation dataset, further comprising: The dataset consisting of the questionnaire data, quantitative CT data, and structured sign data corresponding to the subjects of the derivation cohort was randomly divided into a training cohort and an internal validation cohort in a ratio of 80% and 20% using the createDataPartition function of the caret R package; Dividing the data of the training cohort and the internal validation cohort into measurement data and counting data, respectively, and further dividing the counting data into binary data and ordered multi-classification data; Determine the measurement data as the first sequence; The values ​​of each binary count data are recoded as 1 and 2 to obtain the second series; Re-encoding the ordered multi-classification data by adopting a fusion weight encoding method to obtain a third sequence; Use the missforest R package to interpolate missing values ​​for the first sequence, the second sequence, and the third sequence in the training cohort and the internal validation cohort, respectively, to obtain interpolated data sequences; Using the nearZeroVar function of the caret R package, identifying and removing features of zero variance in the interpolated data sequence corresponding to the training cohort, to obtain a removed data sequence; Use the createFolds function to randomly split the eliminated data sequence corresponding to the training queue into multiple groups, and finally obtain a preprocessed training data set and a preprocessed internal validation data set; The formula for the fusion weight encoding is: ;in, is the maximum integer value after recoding, is the total number of categories, For category For the target variable The contribution weight of ;in, is an event of category i, is the probability that category i and target variable Y occur simultaneously, and They are and the marginal probability of Y, is the value of category i, and y is the specific value of the target variable Y.

2. The method for constructing a COPD screening model based on quantitative CT according to claim 1, characterized in that: The optimal model is constructed based on the XGBoost algorithm.

3. The method for constructing a COPD screening model based on quantitative CT according to claim 1, characterized in that: The study population was screened according to pre-specified criteria to obtain the derivation cohort and the first external validation cohort, including: The subjects from the derivation cohort who met the pre-set inclusion and exclusion criteria were randomly divided into a training cohort and an internal validation cohort at a ratio of 8:2; The preset hospital cohorts were used as the first independent external validation cohorts.

4. The method for constructing a COPD screening model based on quantitative CT according to claim 1, characterized in that: The single-modal modeling includes: modeling using the questionnaire data in the training data set, modeling using the quantitative CT data in the training data set, and modeling using the structured sign data in the training data set; the multi-modal modeling includes: modeling using the questionnaire data and quantitative CT data in the training data set, modeling using the questionnaire data and structured sign data in the training data set, modeling using the quantitative CT data and structured sign data in the training data set, and modeling using the questionnaire data, quantitative CT data, and structured sign data in the training data set.

5. The method for constructing a COPD screening model based on quantitative CT according to claim 3, characterized in that: The internal validation data set was used to validate and evaluate the model after modeling, and the indicators included: AUC, sensitivity, specificity, accuracy, positive predictive value, negative predictive value and F1 score.

Citation Information

Patent Citations

  • Chronic obstructive pulmonary disease predicting system based on migrating study, equipment and medium

    CN111248913A

  • Method for constructing chronic obstructive pulmonary disease early-stage model based on chest CT (Computed Tomography) parameters

    CN118136254A